18 citations · 20 across the 3 of their papers we have counts for
4 papers · 1 filter
Life After Benchmark Saturation: A Case Study of CORE-Bench
Nitya Nadgir, Sayash Kapoor, Kangheng Liu +11
When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the opportunity t…
Towards a Science of AI Agent Reliability
Stephan Rabanser, Sayash Kapoor, Peter Kirgis +3
AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in pr…
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
Sayash Kapoor, Benedikt Stroebl, Peter Kirgis +28
AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of…
MASAI: Modular Architecture for Software-engineering AI Agents
Daman Arora, Atharv Sonwane, Nalin Wadhwa +5
A common method to solve complex problems in software engineering, is to divide the problem into multiple sub-problems. Inspired by this, we propose a Modular Architecture for Soft…