works on

From the 1 of 24 linked papers with an AI index.

most citedCORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

4 citations · 4 across the 5 of their papers we have counts for

collaborators

24 papers

cs.AI2026

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21

The paper evaluates whether current AI agents can independently conduct open‑ended AI research by having them attempt to solve the central questions of two unpublished NeurIPS subm…

cs.CY2026

FLARE-AI: Flaw Reporting for AI

Shayne Longpre, Elaine Zhu, Carson Ezell +15

Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety. Yet the AI reporting ecosystem is fragmented: researchers who identify…

cs.CY2026

Bridging Predictions and Interventions: An Integrated Framework for Automated Decision-Systems

Inioluwa Deborah Raji, Lydia T. Liu, Angela Zhou +27

Automated decision systems (ADS) leverage predictions about individual future outcomes to inform consequential decision-making in organizational settings. Across various settings -…

cs.AI2026

Life After Benchmark Saturation: A Case Study of CORE-Bench

Nitya Nadgir, Sayash Kapoor, Kangheng Liu +11

When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the opportunity t…

cs.CL20264 cited

CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

Zachary S. Siegel, Sayash Kapoor, Nitya Nadgir +2

AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research. To spur the development of useful agents, we need benchmark…

cs.AI2026

Towards a Science of AI Agent Reliability

Stephan Rabanser, Sayash Kapoor, Peter Kirgis +3

AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in pr…