most citedCORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

4 citations · 4 across the 2 of their papers we have counts for

collaborators

7 papers

cs.CY2026

Safety Degradation in AI Agents

Cheng Yu, Benedikt Stroebl, Diyi Yang +1

Despite the growing integration of retrieval-enabled AI agents into society, their safety and ethical behavior remain inadequately understood. In particular, the integration of LLM…

cs.CL20264 cited

CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

Zachary S. Siegel, Sayash Kapoor, Nitya Nadgir +2

AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research. To spur the development of useful agents, we need benchmark…

cs.LG2026

The Limits of Inference Scaling Through Resampling

Benedikt Stroebl, Sayash Kapoor, Arvind Narayanan

Recent research has generated hope that inference scaling, such as resampling solutions until they pass verifiers like unit tests, could allow weaker models to match stronger ones.…

cs.CR2025

Dynamic Risk Assessments for Offensive Cybersecurity Agents

Boyi Wei, Benedikt Stroebl, Jiacen Xu +3

Foundation models are increasingly becoming better autonomous programmers, raising the prospect that they could also automate dangerous offensive cyber-operations. Current frontier…

cs.AI2025

Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

Sayash Kapoor, Benedikt Stroebl, Peter Kirgis +28

AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of…

cs.CL2025

The Book of Life approach: Enabling richness and scale for life course research

Mark D. Verhagen, Benedikt Stroebl, Tiffany Liu +2

For over a century, life course researchers have faced a choice between two dominant methodological approaches: qualitative methods that analyze rich data but are constrained to sm…