arxiv
PublishedJuly 29, 2026 at 4:00 AM
▲bullish
FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon
Publisher summary· verbatim
arXiv:2607.22785v1 Announce Type: cross Abstract: Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units. Prior attempts, including our own, failed or pro
Models mentioned
01Related
05- arxivJul 18Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning
- arxivJul 16A Shared Subcircuit Lets LLMs Count Down Across Tasks
- techcrunchJun 30Vibe-coding platform Base44 launches own model as AI startups seek defensibility
- arxivMay 16Krause Synchronization Transformers
- arxivMay 15A Large Language Model Based Pipeline for Review of Systems Entity Recognition from Clinical Notes
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivCulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning9harxivProactive Service Agents: A Unified Decision Framework, Methods, and Evaluation9harxivX-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System9harxivA Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors9hThe Bubble Brief
WEEKLYRead hardware-acceleration insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗