agent evaluation 1asynchronous runtime 1benchmarking 1large language models 1software infrastructure 1
From the 1 of 5 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
Jize Wang, Xuanxuan Liu, Yining Li +7
The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows. However, current tool-use be…
cs.CL2025
ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
Hongwei Liu, Junnan Liu, Shudong Liu +33
The rapid advancement of Large Language Models (LLMs) has led to performance saturation on many established benchmarks, questioning their ability to distinguish frontier models. Co…