4 papers · 1 filter
A Unified Framework for the Evaluation of LLM Agentic Capabilities
Pengyu Zhu, Lijun Li, Yaxing Lyu +8
As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores often jointly reflect model…
From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control
Rui Ha, Rui Pu, Chaozhuo Li +2
Large Reasoning Models (LRMs) can exhibit step-by-step reasoning, reflection, and backtracking, but these behaviors are often unregulated, leading to overthinking. As a result, LRM…
"LLM Agent Performance" Is Not a Single Evaluation Target
Pengyu Zhu, Li Sun, Philip S. Yu +1
LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Unified execution controls these non-model…
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
Yi Liu, TingFeng Hui, Wei Zhang +4
Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, bri…