1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Zhuochun Li, Youngmin Ko, Ali Keramati +9
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential…