arxiv
PublishedSeptember 14, 2026 at 4:00 AM
—neutral
Dissecting GPU Utilization for LLM Inference on Nvidia Hopper
Publisher summary· verbatim
arXiv:2609.12923v1 Announce Type: cross Abstract: A single SM utilization percentage can make an LLM inference workload look compute-saturated while hiding how much useful work is being done. The problem is not that the counter is wrong, but that it collapses several different mechanisms into one nu
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivBenford's Law as a Distributional Prior for Post-Training Quantization of Large Language Models6harxivTransformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling6harxivEvidence-Aligned Local Composition of Discrete Experts for Sequence Restoration6harxivDynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model6hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗