·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
US threatens sanctions against Chinese AI models over IP theft36m◆Google launches a cheaper alternative to large AI security models like Mythos1h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated2h◆Halliday’s latest smart glasses feature a much-improved display3h◆America needs to stop getting shocked by Chinese AI5h◆Advancing next-gen AI with materials science innovation5h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else6h◆Capacity and Redundancy Trade-offs in Multi-Task Learning12h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation12h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making12h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection12h◆Supervised Reward Inference12h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization12h◆Is Progressive Disclosure All You Need for Long-Context Agents?12h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability12h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification12h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration12h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI12h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models12h◆AI-Augmented Human Resource Management? Insights from German companies12h◆US threatens sanctions against Chinese AI models over IP theft36m◆Google launches a cheaper alternative to large AI security models like Mythos1h◆Music streamer Deezer says more than 50% of daily uploads are AI-generated2h◆Halliday’s latest smart glasses feature a much-improved display3h◆America needs to stop getting shocked by Chinese AI5h◆Advancing next-gen AI with materials science innovation5h◆Gritt exits stealth with $32 million for robots to build solar plants — then, everything else6h◆Capacity and Redundancy Trade-offs in Multi-Task Learning12h◆Predictive Training with Latent Imagination for Visual Quadruped Navigation12h◆Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making12h◆Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection12h◆Supervised Reward Inference12h◆PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization12h◆Is Progressive Disclosure All You Need for Long-Context Agents?12h◆It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability12h◆DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification12h◆Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration12h◆Scientific reasoning does not reliably translate into scientific forecasting in frontier AI12h◆Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models12h◆AI-Augmented Human Resource Management? Insights from German companies12h◆
News/model/privacy-filter

privacy-filter news

47 articles mentioning privacy-filter

arxiv3d ago

Auditing Fairness-Privacy Trade-offs: Subpopulation-Level Effects of Fairness-Enhancing Algorithms

arXiv:2607.14607v1 Announce Type: cross Abstract: Machine learning (ML) models deployed in sensitive domains such as healthcare, law enforcement, and finance must satisfy not only utility requirements but also fairness and privacy guarantees. While prior work has largely examined how privacy-preserv

arxiv3d ago

Privacy Leakage in Federated Learning in Radiology Reports: A Comparative Evaluation of Tokenizer-Driven Privacy Risks

arXiv:2607.14205v1 Announce Type: new Abstract: Federated learning (FL) enables multi-institutional training on clinical text without sharing raw data, but gradient inversion can reconstruct sensitive information from shared model updates. The extent of this leakage for radiology reports, and the ro

arxiv5d ago

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models

arXiv:2607.13093v1 Announce Type: cross Abstract: On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers strong computing power but exposes user prompts and dialogue data, while standalone on-device inference is unfeas

arxiv5d ago

Privacy Preserving Recommender Systems Balancing Personalization with Privacy

arXiv:2607.13328v1 Announce Type: cross Abstract: Personalized recommendation systems are central to modern e-commerce and retail platforms, but they typically rely on centralized storage of detailed user interaction data, creating significant privacy and regulatory challenges. With increasing requi

arxiv5d ago

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge

arXiv:2607.13088v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly moving from research settings into the wild, deployed on enterprise infrastructure, personal devices, and edge platforms. While cloud deployments offer scalable compute, concerns over data sovereignty, complia

arxiv5d ago

What Your Model Threw Away and Why You'll Want It Back: Masking, Fingerprinting, and Privacy from Discarded Geometry

arXiv:2607.13046v1 Announce Type: new Abstract: We develop a framework for the information discarded by machine learning models whose inputs carry a Lie group action. Given a representation $\pi$ of a Lie group $G$ on a space $V$ and a learned function $f\colon V \to \mathbb{R}$, we define two objec

arxiv5d ago

When T2I Synthetic Data Backfires: Amplified Privacy Risks in Real-Synthetic Mix Training

arXiv:2607.13541v1 Announce Type: cross Abstract: To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-

arxiv6d ago

Reducing information dependency does not cause training data privacy. Adversarially non-robust features do

arXiv:2607.12354v1 Announce Type: new Abstract: In this paper, we challenge the prevailing view that information dependency (including rote memorization) drives training data exposure to image reconstruction attacks. We show that extensive exposure can persist without rote memorization and is instea

arxiv6d ago

Toward Production-Ready Federated Learning in Healthcare: Privacy, Orchestration, and Governance in MLOps

arXiv:2607.10467v2 Announce Type: replace-cross Abstract: Healthcare organizations often cannot freely centralize patient data because medical records are sensitive, regulated, and institutionally controlled. Federated learning offers a practical alternative by allowing hospitals and clinics to trai

arxiv6d ago

Proximity Features: Privacy-Compliant Cold-Start Personalization at Airbnb

arXiv:2607.12246v1 Announce Type: new Abstract: Personalization in two-sided marketplaces relies heavily on user-level features, yet for platforms with infrequent, high-consideration purchases, a large fraction of users lack sufficient history for effective recommendation, spanning both paid and org

arxiv6d ago

Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model

arXiv:2606.26373v4 Announce Type: replace-cross Abstract: Semantic search creates an asymmetric disclosure problem: query embeddings may reveal user intent, while returning exact provider vectors distributes reusable representations. We evaluate a deliberately restricted hybrid design. A public corp

arxivJul 14

IntraShuffler: A Privacy Preserving Framework for Heterogeneous DP Federated Learning

arXiv:2606.02563v2 Announce Type: replace Abstract: Heterogeneous Differential Privacy (HDP) in Federated Learning (FL) allows clients to select individual privacy budgets ($\varepsilon_i$) according to institutional policies and data sensitivity. In practice, many HDP-FL systems employ $\varepsilon

arxivJul 14

Privacy-Aware Collaborative and Distributed Bayesian Optimization

arXiv:2607.11600v1 Announce Type: new Abstract: We propose a collaborative meta-learning framework for distributed Bayesian optimization matching centralized performance without raw-data exchange. We show gradient sharing leaks client observations, with leakage worsening as the search converges and

arxivJul 14

PromptGraph: Graph-Guided Prompt Sanitization for Balancing Privacy and Utility in LLM Inference

arXiv:2607.10709v1 Announce Type: cross Abstract: Large Language Model (LLM) services introduce a fundamental privacy challenge. Sensitive information may be inferred not only from explicit identifiers, such as names or phone numbers, but also from contextual associations among otherwise innocuous s

arxivJul 14

Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation

arXiv:2604.07486v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing synthetic replicas of private text, which requires carefully balancing privacy and utility. We propose

arxivJul 13

How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy

arXiv:2512.03238v2 Announce Type: replace-cross Abstract: High quality data is needed to unlock the full potential of AI for end users. However finding new sources of such data is getting harder: most publicly-available human generated data will soon have been used. Additionally, publicly available

arxivJul 13

LDPKiT: Superimposing Remote Queries for Privacy-Preserving Distillation

arXiv:2405.16361v4 Announce Type: replace Abstract: To protect privacy in regulated domains such as healthcare and finance, model owners may allow only remote API access while keeping both the training data and model parameters private. However, model users performing inference on such remotely host

arxivJul 11

Measuring the practice of shared-decision making (OPTION12): An Investigation into Open-sourced Smaller LLMs (OS-sLLMs) for Better Privacy and Sustainability

arXiv:2607.06127v2 Announce Type: replace Abstract: We present LLM4SDM, the first study of open-source smaller language models (OS-sLLMs) for automated assessment of shared decision making (SDM) using the Observer OPTION12 framework. Unlike previous work that relies on large commercial models and th

arxivJul 10

Swapping Faces, Saving Features: A Dual-Purpose Pipeline for Pedestrian Privacy in ITS

arXiv:2607.08402v1 Announce Type: cross Abstract: Large-scale and diverse datasets are needed to train AI models to take real-time decisions for autonomous vehicles (AVs), an intelligent transportation system (ITS) application. Pedestrian intention and trajectory prediction are critical models used

arxivJul 10

Idiobionics: The Unification of Privacy and Intelligent Robotic Prostheses

arXiv:2607.07775v1 Announce Type: new Abstract: The human body is at the center of a growing family of technologies designed to tightly and persistently couple biological and digital systems. Robotic prostheses are a representative example of this tight coupling. Also referred to as bionic limbs, ro

#robotics#prosthetics#security
arxivJul 10

EdgeRefine: Privacy-Utility Balance for Graphs via Jaccard Sampling under Edge Differential Privacy

arXiv:2607.08659v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) have shown considerable success in learning from graph-structured data, but their use in privacy-sensitive areas remains difficult because graph structure can leak sensitive link information. To satisfy edge-level different

arxivJul 10

Federated Deep Learning for Privacy-Preserving Cardiovascular Disease Risk Prediction

arXiv:2607.08595v1 Announce Type: new Abstract: Cardiovascular disease risk prediction models often rely on data from a single institution or centrally pooled datasets. Extending these models across institutions could be limited by privacy regulations and constraints on sharing patient-level data. F

arxivJul 10

Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models

arXiv:2607.08282v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have become essential productivity tools, their integration into workflows without adequate safeguards creates significant risks. This paper proposes an open-source, privacy-focused, user-facing firewall designed to

arxivJul 3

Split-n-Chain: Privacy-Preserving Multi-Node Split Learning with Blockchain-Based Auditability

arXiv:2503.07570v3 Announce Type: replace-cross Abstract: Deep learning, when integrated with a large amount of training data, has the potential to outperform machine learning in terms of high accuracy. Recently, privacy-preserving deep learning has drawn significant attention of the research commun

arxivJul 3

Privacy-Preserving and Verifiable Approximate Distributed Coded Computing

arXiv:2607.02187v1 Announce Type: new Abstract: Distributed machine learning enables collaborative model training without centralizing data, but it also exposes learning processes to privacy leakage and malicious manipulation. Existing defenses typically address these threats in isolation and are of

arxivJul 3

Unveiling the Non-Monotonic Effect of Privacy on Generalization under Byzantine Robustness

arXiv:2607.01492v1 Announce Type: new Abstract: Recent work has established a fundamental trilemma between Byzantine robustness, local differential privacy (LDP), and optimization error in distributed learning. We show that this trilemma does not universally extend to generalization error, but inste

arxivJul 3

AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs

arXiv:2601.22710v2 Announce Type: replace-cross Abstract: Modern LLMs are increasingly accessed via black-box APIs, requiring users to transmit sensitive prompts, outputs, and fine-tuning data to external providers, creating a critical privacy risk at the API boundary. We introduce AlienLM, a deploy

arxivJul 2

NeuroFilter: Activation-Based Guardrails for Privacy-Conscious LLM Agents

arXiv:2601.14660v2 Announce Type: replace-cross Abstract: Agentic Large Language Models (LLMs) are models able to reason, plan, and execute tools over unstructured data. These abilities are enabling transformative applications in domains spanning from personal assistant, financial, and legal domains

#privacy#security#language-models
techcrunchJul 1

Venice AI becomes a unicorn with $65M Series A as its privacy-first AI platform takes off

Venice AI is already profitable, with annualized run-rate revenues of over $70 million, CEO Erik Voorhees said.

arxivJul 1

Semantic Leakage and Privacy Preservation in Relay-Assisted Semantic Communications

arXiv:2606.31973v1 Announce Type: cross Abstract: Semantic communication (SemCom) has emerged as a promising paradigm in which the transmission of task-relevant information is prioritized over raw data, enabling efficient and robust communication under resource and channel constraints. In this paper

techcrunchJun 30

Lumo, Proton’s privacy-focused AI chatbot, gets an upgrade

Proton's Lumo 2.0 is dropping this week, giving users a broader variety of capabilities.

arxivJun 30

Efficient Unlearning with Privacy Guarantees

arXiv:2507.04771v2 Announce Type: replace-cross Abstract: Privacy protection laws, such as the GDPR, grant individuals the right to request the forgetting of their personal data not only from databases but also from machine learning (ML) models trained on them. Machine unlearning has emerged as a pr

arxivJun 30

Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict

arXiv:2601.03546v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate decision-making tasks involving personal data sharing, where privacy concerns and prosocial motivations can push choices in opposite directions. Existing evaluations often measure

arxivJun 30

SurrogateShield: Beyond Redaction for High-Utility, Privacy-Preserving LLM Interactions

arXiv:2606.29567v1 Announce Type: cross Abstract: LLM-based assistants transmit user queries verbatim to third-party API endpoints that lie outside the user's audit or control. When those queries contain personally identifiable information (PII), the data persists on remote infrastructure subject to

arxivJun 30

Decomposing Memorization Reduction in Privacy-Preserving Fine-Tuning of SLMs for CSIRTs

arXiv:2606.28479v1 Announce Type: cross Abstract: CSIRTs increasingly fine tune language models on vulnerability scan records, but these records expose internal network topology and create privacy risks under regulations such as GDPR and LGPD. We present the first empirical study of how DP SGD and H

arxivJun 29

Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy

arXiv:2606.27936v1 Announce Type: cross Abstract: The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by the public. While prior research has established that mobility traces are highly unique and that indi

arxivJun 29

Productionized Fairness Measurement Under Privacy Constraints

arXiv:2606.27558v1 Announce Type: new Abstract: Fairness measurements in the form of disaggregated evaluations often rely on demographic signals that are legally constrained or culturally sensitive. Race and ethnicity signals are among the more difficult signals to curate and use for this task. This

arxivJun 29

ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents

arXiv:2606.28061v1 Announce Type: cross Abstract: Large language models (LLMs) have increasingly moved from standalone text generation systems to agents that invoke external tools, access environments, and execute multi-step tasks. However, conventional function-calling benchmarks mainly evaluate ta

arxivJun 27

TGHE: Template-based Graph Homomorphic Encryption for Privacy-Preserving GNN Inference in Edge-Cloud Systems

arXiv:2606.26664v1 Announce Type: cross Abstract: Existing homomorphic encryption (HE)-based GNN systems adopt a graph-centric paradigm that couples per-query cost to global graph size, limiting evaluations to at most ~20k nodes and making them incompatible with dynamic, large-scale financial graphs

arxivJun 27

Privacy-Aware Agent Collaboration for Dynamic VR Slice Management in 6G SD-RAN

arXiv:2606.26123v1 Announce Type: cross Abstract: Ultra-low latency and high throughput are required for Virtual Reality (VR) services in 6G networks, which presents critical challenges for Software-Defined Radio Access Networks (SD-RANs) dynamic resource management. This work propose a mobility-dri

arxivJun 27

Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents

arXiv:2606.26627v1 Announce Type: cross Abstract: Large language model agents increasingly query databases, search document collections, call external APIs, remember past interactions, and act on a user's behalf. As they move from answering questions to operating over sensitive data, privacy becomes

arxivJun 26

ProfileFoundry: A Synthetic Person-Object Substrate for Privacy, Memory, and Tool-Use Evaluation in LLM Agent

arXiv:2606.26403v1 Announce Type: new Abstract: Foundation-model research increasingly needs data about people: user state, personal histories, relationships, contact-like fields, documents, and longitudinal updates. Real user data is difficult to share, perturb, audit, or redistribute responsibly,

arxivJun 25

Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity

arXiv:2606.24623v1 Announce Type: cross Abstract: Retrieval-Augmented Generation enhances large language models by incorporating external knowledge, but deploying it in sensitive scenarios risks privacy leakage via malicious prompts. To address this, we propose a multi-agent framework that sanitizes

arxivJun 25

Security and Privacy in Retrieval-Augmented Generation: Architectures, Threats, Defenses, and Future Directions for Building Trustworthy Systems

arXiv:2606.25533v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has emerged as a dominant paradigm for enhancing large language models with external knowledge. By coupling retrieval mechanisms with generative models, RAG systems improve factual grounding and adaptability acros

arxivJun 25

Privacy-Aware Visual Language Models

arXiv:2405.17423v4 Announce Type: replace-cross Abstract: As Visual Language Models (VLMs) become increasingly embedded in everyday applications, ensuring they can recognise and appropriately handle privacy-sensitive content is thus essential to protect users. To this end, we conduct a comprehensive

arxivJun 25

TL++: Accuracy and Privacy Preserving Traversal Learning for Distributed Intelligent Systems

arXiv:2606.25627v1 Announce Type: new Abstract: Distributed intelligent systems increasingly need to train across data silos without centralizing raw data. Federated learning keeps data local but can suffer under heterogeneous partitions and requires repeated full-model exchange. Split learning redu

arxivJun 24

Natural Identifiers for Privacy and Data Audits in Large Language Models

arXiv:2606.24408v1 Announce Type: new Abstract: Assessing the privacy of large language models (LLMs) presents significant challenges. In particular, most existing methods for auditing differential privacy require the insertion of specially crafted canary data during training, making them impractica

HomeModelsNews