6 papers
The Foreign Policy AI Evaluation Gap
Charles Pozniak, Jeba Sania
We argue that AI systems used in conducting foreign policy tasks - broadly enacting 'statecraft' - should be a priority test case for technical AI governance research. In enacting…
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Mubashara Akhtar, Anka Reuel, Prajna Soni +36
Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly "saturate", making it difficult…
From Democracies to Autocracies: How AI Systems Enable Authoritarianism by Design
Jeba Sania, Marta Ziosi, Fazl Barez
AI-enabled authoritarianism is not confined to autocracies. In this paper, we provide greater transparency by investigating and mapping the lifecycles of six AI systems deployed in…
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
Anka Reuel, Avijit Ghosh, Jenny Chim +32
Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their risks and capabilities. Although general c…
Expert Survey: AI Reliability & Security Research Priorities
Joe O'Brien, Jeremy Dolan, Jay Kim +5
Our survey of 53 specialists across 105 AI reliability and security research areas identifies the most promising research prospects to guide strategic AI R&D investment. As compani…