arxiv
PublishedJuly 20, 2026 at 4:00 AM
—neutral
Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length
Publisher summary· verbatim
arXiv:2512.07019v3 Announce Type: replace-cross Abstract: The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream applications and actionable future improvements. The Item Response Theory (IRT) model with Computerized Adaptive T
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivADS-C: Antidistillation Sampling for Classification20harxivTesting Distributions Against Bounded Distinguishers20harxivBeyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes20harxivFrom Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems20hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗