From the 1 of 5 linked papers with an AI index.
5 papers
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
Kai Chen, Zichen Ding, Jiaye Ge +22
The paper presents AgentCompass, an open‑source infrastructure that standardizes and simplifies the evaluation of large‑language‑model based autonomous agents by separating benchma…
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
Jize Wang, Xuanxuan Liu, Yining Li +7
The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows. However, current tool-use be…
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
Yiheng Wang, Yixin Chen, Shuo Li +33
We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike gene…
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
Mingqi Wu, Zhihao Zhang, Qiaole Dong +11
Reasoning in large language models has long been a central research focus, and recent studies employing reinforcement learning (RL) have introduced diverse methods that yield subst…
ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
Hongwei Liu, Junnan Liu, Shudong Liu +33
The rapid advancement of Large Language Models (LLMs) has led to performance saturation on many established benchmarks, questioning their ability to distinguish frontier models. Co…