Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
Keyu Li, Junhao Shi, Yang Xiao +11
Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. However, existing benchmarks remain f…
cs.AI2026
Credit-Budgeted ICPC-Style Coding: When Agents Must Pay for Every Decision
Lingfeng Zhou, Junhao Shi, Jin Gao +1
Current evaluations of autonomous coding agents assume an unrealistic, infinite-resource environment. However, real-world software engineering is a resource-bound competition. As w…