Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards
Keyu Li, Jin Gao, Dequan Wang
On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that fa…
cs.AI2026
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
Keyu Li, Junhao Shi, Yang Xiao +11
Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. However, existing benchmarks remain f…
cs.AI2026
Credit-Budgeted ICPC-Style Coding: When Agents Must Pay for Every Decision
Lingfeng Zhou, Junhao Shi, Jin Gao +1
Current evaluations of autonomous coding agents assume an unrealistic, infinite-resource environment. However, real-world software engineering is a resource-bound competition. As w…