arxiv
PublishedApril 10, 2026 at 4:00 AM
—neutral
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
Publisher summary· verbatim
arXiv:2603.28281v2 Announce Type: replace Abstract: We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset $D$ of trajectory-preference tuples (each preference being an $n$-dimensio
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivGaussian Linear Functional Manifold Method for Massive Point Cloud Data9harxivMind the Gap: Navigating Inference with Optimal Transport Maps9harxivWhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Small Pretrained Speech Recognition Transformers9harxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9hThe Bubble Brief
WEEKLYRead machine-learning insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗