1 paper
Jiaqi Shao, Hanck Chen, Wei Zhang +2
Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation…