1 paper · 1 filter
Amanda Dsouza, Ramya Ramakrishnan, Charles Dickens +2
As AI agents integrate into enterprise applications, their evaluation demands benchmarks that reflect the complexity of real-world operations. Instead, existing benchmarks overemph…