·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear1h◆What’s at stake in AI’s trillion-dollar gamble3h◆JaxAHT: A JAX-Based Library for Ad Hoc Teamwork9h◆MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents9h◆Signatures of Steerability in Activation Space of Language Models9h◆Data-free On-policy Distillation9h◆A latent dimension of Condorcet's jury theorem for multiple AI advisers9h◆Personalizing Personal Health Interfaces: Co-Design with Generative AI9h◆Ensemble Complexity in Photovoltaic Forecasting9h◆Self-Evolving Memory for Generative Recommendation9h◆Mecha-nudges for Machines9h◆Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations9h◆Mimir: Large-scale Multilingual Concept Modeling9h◆Toward a Layer-2 Trigger for AI/ML Lifecycle Management in 6G9h◆How User-AI Mistreatment Occurs and Matters in Conversational Systems?9h◆Solar Intelligence9h◆Prefix Sharing Is a Sorting Problem9h◆Enabling Creative Exploration for Vibe Design Agents9h◆Recurrent GraphNeural NetworkswithSet-BasedAggregation9h◆Attention Is All You Need (to Avoid Spurious Oscillations)9h◆Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear1h◆What’s at stake in AI’s trillion-dollar gamble3h◆JaxAHT: A JAX-Based Library for Ad Hoc Teamwork9h◆MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents9h◆Signatures of Steerability in Activation Space of Language Models9h◆Data-free On-policy Distillation9h◆A latent dimension of Condorcet's jury theorem for multiple AI advisers9h◆Personalizing Personal Health Interfaces: Co-Design with Generative AI9h◆Ensemble Complexity in Photovoltaic Forecasting9h◆Self-Evolving Memory for Generative Recommendation9h◆Mecha-nudges for Machines9h◆Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations9h◆Mimir: Large-scale Multilingual Concept Modeling9h◆Toward a Layer-2 Trigger for AI/ML Lifecycle Management in 6G9h◆How User-AI Mistreatment Occurs and Matters in Conversational Systems?9h◆Solar Intelligence9h◆Prefix Sharing Is a Sorting Problem9h◆Enabling Creative Exploration for Vibe Design Agents9h◆Recurrent GraphNeural NetworkswithSet-BasedAggregation9h◆Attention Is All You Need (to Avoid Spurious Oscillations)9h◆
News/Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
arxiv
PublishedSeptember 15, 2026 at 4:00 AM
—neutral

Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2604.27093v2 Announce Type: replace-cross Abstract: Current LLM safety alignment techniques improve model robustness against adversarial attacks, but overlook whether and how LLMs can recover helpfulness when benign users clarify their intent. We introduce CarryOnBench, the first interactive b

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →

Related coverage

More from ARXIV
arxivJaxAHT: A JAX-Based Library for Ad Hoc Teamwork9harxivMOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents9harxivSignatures of Steerability in Activation Space of Language Models9harxivData-free On-policy Distillation9h
The Bubble Brief
WEEKLY

Read AI insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews