From the 3 of 12 linked papers with an AI index.
9 papers · 1 filter
Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks
Jeff Mohl, Nelson Gardner-Challis, Magda Dubois +6
The paper presents automated AI scanners that analyze benchmark transcripts to detect validity flaws such as ground‑truth leakage, tool failures, guessing vulnerabilities, and ambi…
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21
The paper evaluates whether current AI agents can independently conduct open‑ended AI research by having them attempt to solve the central questions of two unpublished NeurIPS subm…
How Inference Compute Shapes Frontier LLM Evaluation
Jessica McFadyen, Ole Jorgensen, Harry Coppock +2
The paper studies how the amount of compute allocated during inference (e.g., token budget, repeated attempts) affects the performance of frontier large language models on challeng…
Open-World Evaluations for Measuring Frontier AI Capabilities
Sayash Kapoor, Peter Kirgis, Andrew Schwartz +15
Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it privileges tasks that can be…
Seven simple steps for log analysis in AI systems
Magda Dubois, Ekin Zorer, Maia Hamin +17
AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensities, and behaviours, or assess…
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
Ee Wei Seah, Yongsen Zheng, Naga Nikshith +67
The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing rema…