Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
Manik Rana, Calissa Man, Anotida Expected Msiiwa +5
Goal changes are a defining feature of real world multi-turn interactions, yet current agent benchmarks primarily evaluate static objectives or one-shot tool use. We introduce Agen…
cs.AI2025
Limits of Emergent Reasoning of Large Language Models in Agentic Frameworks for Deterministic Games
Chris Su, Harrison Li, Matheus Marques +3
Recent work reports that Large Reasoning Models (LRMs) undergo a collapse in performance on solving puzzles beyond certain perplexity thresholds. In subsequent discourse, questions…