1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Deepak Akkil, Mowafak Allaham, Amal Raj +2
Reliable evaluation of AI agents operating in complex, real-world environments requires methodologies that are robust, transparent, and contextually aligned with the tasks agents a…