Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise
Abdulrahman AlRabah, Xiaocheng Yang, Dilek Hakkani-Tür +1
Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains such as education, healthcar…
cs.CL2025
EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents
Cheng Qian, Peixuan Han, Qinyu Luo +9
Language model agents excel in long-session planning and reasoning, but existing benchmarks primarily focus on goal-oriented tasks with explicit objectives, neglecting creative ada…