arxiv
PublishedJune 2, 2026 at 4:00 AM
▲bullish
AdaCodec: A Predictive Visual Code for Video MLLMs
Publisher summary· verbatim
arXiv:2606.02569v1 Announce Type: cross Abstract: Video is temporally redundant: adjacent frames usually share most objects, background, and layout. Yet existing video multimodal large language models (video MLLMs) usually encode each sampled frame as an independent RGB image, causing visual tokens
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivSFMambaNet: Spectral-Frequency Enhanced Selective State Space Model for Correspondence Pruning13harxivOptical-Guided Neural Collapse for SAR Few-Shot Class Incremental Learning13harxivDynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models13harxivTemporal Order Matters for Agentic Memory: Segment Trees for Long-Horizon Agents13hThe Bubble Brief
WEEKLYRead video insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗