6 papers · 1 filter
The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards
Keyu Li, Jin Gao, Dequan Wang
On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that fa…
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
Keyu Li, Junhao Shi, Yang Xiao +11
Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. However, existing benchmarks remain f…
Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training
Dayuan Fu, Yunze Wu, Xiaojie Cai +13
Large Language Model (LLM) agents have recently shown strong potential in domains such as automated coding, deep research, and graphical user interface manipulation. However, train…
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
Yunze Wu, Dayuan Fu, Weiye Si +13
AI agents could accelerate scientific discovery by automating hypothesis formation, experiment design, coding, execution, and analysis, yet existing benchmarks probe narrow skills…
LIMI: Less is More for Agency
Yang Xiao, Mohan Jiang, Jie Sun +18
We define Agency as the emergent capacity of AI systems to function as autonomous agents actively discovering problems, formulating hypotheses, and executing solutions through self…
DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery
Keyu Li, Mohan Jiang, Dayuan Fu +4
The rapid advancement of large language models has fundamentally shifted the bottleneck in AI development from computational power to data availability-with countless valuable data…