1 paper
Amanda Dsouza, Ramya Ramakrishnan, Charles Dickens +2
As AI agents integrate into enterprise applications, their evaluation demands benchmarks that reflect the complexity of real-world operations. Instead, existing benchmarks overemph…