1 paper
Mina Mohammadmirzaei, Jeffrey Flanigan
Computer-use agents are increasingly evaluated by whether they complete realistic desktop and web tasks. However, task success alone can miss failures in which an agent reaches the…