works on

From the 2 of 12 linked papers with an AI index.

collaborators

12 papers

cs.AI2026

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks

Jeff Mohl, Nelson Gardner-Challis, Magda Dubois +6

The paper presents automated AI scanners that analyze benchmark transcripts to detect validity flaws such as ground‑truth leakage, tool failures, guessing vulnerabilities, and ambi…

cs.AI2026

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21

The paper evaluates whether current AI agents can independently conduct open‑ended AI research by having them attempt to solve the central questions of two unpublished NeurIPS subm…

cs.AI2026

Open-World Evaluations for Measuring Frontier AI Capabilities

Sayash Kapoor, Peter Kirgis, Andrew Schwartz +15

Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it privileges tasks that can be…

cs.AI2026

Log analysis is necessary for credible evaluation of AI agents

Peter Kirgis, Sayash Kapoor, Stephan Rabanser +8

Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated or deflated by shortcuts and…

cs.HC2026

Ask don't tell: Reducing sycophancy in large language models

Magda Dubois, Cozmin Ududec, Christopher Summerfield +1

Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-sta…

cs.AI2026

Seven simple steps for log analysis in AI systems

Magda Dubois, Ekin Zorer, Maia Hamin +17

AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensities, and behaviours, or assess…