·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Mathematicians want proof OpenAI didn’t use their work1h◆Powering AI is an architecture problem1h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning8h◆What Does an LLM-Agent Leaderboard Rank Actually Compare?8h◆Aes3D: Aesthetic Assessment in 3D Gaussian Splatting8h◆The Transformer as a Polar State Estimator8h◆WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting8h◆CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation8h◆Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method8h◆Learning Counterfactual World Models for Embodied Reasoning under Partial Observability8h◆Agentic Electronic Design Automation: A Handoff Perspective8h◆Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior8h◆Field-level weak lensing cosmology with $60$ simulations using multifidelity simulation-based inference8h◆Discovering Latent Groups for Robust Classification8h◆Microcanonical Hamiltonian Monte Carlo and the Helmholtz Theorem8h◆TimeWarp: Evaluating Web Agents by Revisiting the Past8h◆What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents8h◆PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding8h◆SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia8h◆Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents8h◆Mathematicians want proof OpenAI didn’t use their work1h◆Powering AI is an architecture problem1h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning8h◆What Does an LLM-Agent Leaderboard Rank Actually Compare?8h◆Aes3D: Aesthetic Assessment in 3D Gaussian Splatting8h◆The Transformer as a Polar State Estimator8h◆WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting8h◆CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation8h◆Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method8h◆Learning Counterfactual World Models for Embodied Reasoning under Partial Observability8h◆Agentic Electronic Design Automation: A Handoff Perspective8h◆Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior8h◆Field-level weak lensing cosmology with $60$ simulations using multifidelity simulation-based inference8h◆Discovering Latent Groups for Robust Classification8h◆Microcanonical Hamiltonian Monte Carlo and the Helmholtz Theorem8h◆TimeWarp: Evaluating Web Agents by Revisiting the Past8h◆What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents8h◆PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding8h◆SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia8h◆Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents8h◆
News/Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning
arxiv
PublishedSeptember 10, 2026 at 4:00 AM

Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2609.10142v1 Announce Type: new Abstract: Large language models remain fragile against malicious fine-tuning, motivating training-time defenses against harmful persona drift. Preventative Steering injects undesirable-trait persona vectors during fine-tuning and removes them at evaluation time,

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning8harxivWhat Does an LLM-Agent Leaderboard Rank Actually Compare?8harxivAes3D: Aesthetic Assessment in 3D Gaussian Splatting8harxivThe Transformer as a Polar State Estimator8h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews