1 paper
Zhuochun Li, Youngmin Ko, Ali Keramati +9
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential…