arxiv
PublishedJune 2, 2026 at 4:00 AM
—neutral
Adaptive Exploration for Latent-State Bandits
Publisher summary· verbatim
arXiv:2602.05139v3 Announce Type: replace Abstract: We study bandits whose rewards depend on an unobserved Markov state that evolves independently of the learner's actions. The optimal arm can change even though the learner observes only past actions and rewards. We propose algorithms that feed LinU
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivBringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning9harxivSubagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks9harxivDistribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts9harxivIn RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning9hThe Bubble Brief
WEEKLYRead bandits insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗