1 paper · 1 filter
Zhiyuan Peng, Xin Yin, Chenhao Ying +5
Existing agent benchmarks primarily test task completion, tool use, or skill utility, but do not isolate whether a runtime can convert evidence from its own runs into reusable skil…