1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Shenghan Zheng, Zonglin Di, Yimin Liu +19
LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from…