2 papers
cs.CL2026
Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding
Eileen Ye, Jiawen Tao, Yaoming Li +7
Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, tracking later revisions, identif…
cs.AI2026
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
Weihuang Zheng, Tianyuan Zou, Eileen Ye +5
Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, a…