1 paper · 1 filter
Zhonghao Zhan, Yefan Zhang, Krinos Li +1
Simulated evaluation is widely used to benchmark AI agents, yet how much evidence a simulated pass provides about physical deployment has not been systematically quantified. We pre…