activity
20172026
most citedThe Effectiveness of Discretization in Forecasting: An Empirical Study on Neural Time Series Models

5 citations · 9 across the 11 of their papers we have counts for

collaborators
Showing 2026Show all

5 papers · 1 filter

cs.AI2026

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations eithe…

cs.AI2026

Life After Benchmark Saturation: A Case Study of CORE-Bench

Nitya Nadgir, Sayash Kapoor, Kangheng Liu +11

When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the opportunity t…

cs.AI2026

Open-World Evaluations for Measuring Frontier AI Capabilities

Sayash Kapoor, Peter Kirgis, Andrew Schwartz +15

Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it privileges tasks that can be…

cs.AI2026

Log analysis is necessary for credible evaluation of AI agents

Peter Kirgis, Sayash Kapoor, Stephan Rabanser +8

Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated or deflated by shortcuts and…

cs.AI20262 cited

Towards a Science of AI Agent Reliability

Stephan Rabanser, Sayash Kapoor, Peter Kirgis +3

AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in pr…