1 paper · 1 filter
Zhi Han, Chenxi Zeng, Liuhaichen Yang +3
LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executi…