1 paper
Tianchen Guan, Xinlei Lin, Royce Cheng-Yue +2
Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an under-instrumented middle layer: realisti…