arxiv
PublishedJuly 27, 2026 at 4:00 AM
—neutral
Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging
Publisher summary· verbatim
arXiv:2607.10428v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) as multi-turn conversational partners requires probing capabilities that single-turn benchmarks miss: persona consistency, evolving intent tracking, emotional dynamics, and goal completion across many turns.
Stay posted· Newsletter
A 5-min weekly brief — top movers, price watch, story of the week.
Discussion
No replies yet. Be first.
Related coverage
More from ARXIV
arxivAgentic Permissions Policy Algebra for Taint Confinement in LLM Agents3harxivBeyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks3harxivThe One-Word Census: Answer-Choice Conformity Across 44 Language Models3harxivCreative Integration: A Decidable Criterion of Creativity3hThe Bubble Brief
WEEKLYRead AI insights every Tuesday — top movers, new releases, story of the week.
Originally published on arxiv ↗