arxiv
PublishedJuly 21, 2026 at 4:00 AM
Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels
Publisher summary· verbatim
arXiv:2607.09796v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) has become an important method for aligning large language models (LLMs) with human preferences because it removes the need for explicit reward modeling and reinforcement learning. However, its performance depen
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivCapacity and Redundancy Trade-offs in Multi-Task Learning12harxivPredictive Training with Latent Imagination for Visual Quadruped Navigation12harxivWhere Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making12harxivDid We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection12hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗