1 paper
Veronica Chatrath, Bryan Zhu, Jingxuan Fan +15
An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, w…