1 paper · 1 filter
Shuaiqi Wang, Aadyaa Maddi, Zinan Lin +1
Today, tool-calling agents are commonly evaluated or tested on static datasets of execution traces, including input commands, agent responses, and associated tool calls. However, i…