papers
Publications (2)
cs.AI2026
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21
The paper evaluates whether current AI agents can independently conduct open‑ended AI research by having them attempt to solve the central questions of two unpublished NeurIPS subm…
#ai agents#research automation#evaluation methodology#failure analysis
cs.CY2023
An International Consortium for Evaluations of Societal-Scale Risks from Advanced AI
Ross Gruetzemacher, Alan Chan, Kevin Frazier +11
Given rapid progress toward advanced AI and risks from frontier AI systems (advanced AI systems pushing the boundaries of the AI capabilities frontier), the creation and implementa…