1 paper · 1 filter
Shrey Nag, Sachita, Abhishek Kumar Singh +2
Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion (AgentBench) or security robustness (AgentDojo, ASB), rather than the complete pipe…