·
DataBubble
  • Home
  • Models
  • News
  • Compare
  • Boards
  • Pricing
  • About
  • Newsletter
  • Methodology
  • Contact
Latest
Mathematicians want proof OpenAI didn’t use their work1h◆Powering AI is an architecture problem2h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9h◆What Does an LLM-Agent Leaderboard Rank Actually Compare?9h◆Aes3D: Aesthetic Assessment in 3D Gaussian Splatting9h◆The Transformer as a Polar State Estimator9h◆WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting9h◆CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation9h◆Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method9h◆Learning Counterfactual World Models for Embodied Reasoning under Partial Observability9h◆Agentic Electronic Design Automation: A Handoff Perspective9h◆Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior9h◆Field-level weak lensing cosmology with $60$ simulations using multifidelity simulation-based inference9h◆Discovering Latent Groups for Robust Classification9h◆Microcanonical Hamiltonian Monte Carlo and the Helmholtz Theorem9h◆TimeWarp: Evaluating Web Agents by Revisiting the Past9h◆What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents9h◆PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding9h◆SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia9h◆Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents9h◆Mathematicians want proof OpenAI didn’t use their work1h◆Powering AI is an architecture problem2h◆Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9h◆What Does an LLM-Agent Leaderboard Rank Actually Compare?9h◆Aes3D: Aesthetic Assessment in 3D Gaussian Splatting9h◆The Transformer as a Polar State Estimator9h◆WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting9h◆CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation9h◆Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method9h◆Learning Counterfactual World Models for Embodied Reasoning under Partial Observability9h◆Agentic Electronic Design Automation: A Handoff Perspective9h◆Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior9h◆Field-level weak lensing cosmology with $60$ simulations using multifidelity simulation-based inference9h◆Discovering Latent Groups for Robust Classification9h◆Microcanonical Hamiltonian Monte Carlo and the Helmholtz Theorem9h◆TimeWarp: Evaluating Web Agents by Revisiting the Past9h◆What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents9h◆PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding9h◆SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia9h◆Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents9h◆
News/Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
arxiv
PublishedApril 10, 2026 at 4:00 AM
—neutral

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback

Source
arxiv.orgfull article ↗
Read on arxiv→
Publisher summary· verbatim

arXiv:2603.28281v2 Announce Type: replace Abstract: We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset $D$ of trajectory-preference tuples (each preference being an $n$-dimensio

Stay posted· Newsletter

A 5-min weekly brief — top movers, price watch, story of the week.

// no spam · unsubscribe one-click · free forever

Discussion
Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#machine-learning#reinforcement-learning#robustness#adversarial-attacks

No replies yet. Be first.

Source
↗
arxiv
Read original ↗All from arxiv →
Tags
04
#machine-learning#reinforcement-learning#robustness#adversarial-attacks

Related coverage

More from ARXIV
arxivGaussian Linear Functional Manifold Method for Massive Point Cloud Data9harxivMind the Gap: Navigating Inference with Optimal Transport Maps9harxivWhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Small Pretrained Speech Recognition Transformers9harxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9h
The Bubble Brief
WEEKLY

Read machine-learning insights every Tuesday — top movers, new releases, story of the week.

// no spam · unsubscribe one-click · free forever

Originally published on arxiv ↗
HomeModelsNews