arxiv
PublishedJuly 18, 2026 at 4:00 AM
—neutral
Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models
Publisher summary· verbatim
arXiv:2607.14552v1 Announce Type: cross Abstract: A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors. When sampling fails, a common fix sh
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivBeyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal6harxivSkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents6harxivPolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment6harxivDo Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?6hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗