2 papers
cs.CL2026
CocoaBench: Evaluating Unified Digital Agents in the Wild
CocoaBench Team, Shibo Hao, Zhining Zhang +29
LLM agents now perform strongly in software engineering, deep research, GUI automation, and various other applications, while recent agent scaffolds and models are increasingly int…
cs.AI2026
SourceBench: Can AI Answers Reference Quality Web Sources?
Hexi Jin, Stephen Liu, Yuheng Li +2
Large language models (LLMs) increasingly answer queries by citing web sources, but existing evaluations emphasize answer correctness rather than evidence quality. We introduce Sou…