1 paper · 1 filter
Tomer Keren, Nitay Calderon, Asaf Yehudai +3
As agent capabilities advance, existing benchmarks, such as I¨2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and lab…