From the 1 of 3 linked papers with an AI index.
Showing cs.CYShow all
2 papers · 1 filter
cs.CY2026
The Benchmark Ceiling: Human Judgment, Evaluation Scarcity, and the Political Economy of AI Capability Measurement
Mark Esposito, Liu Zhang
The paper examines how AI benchmarks lose discriminative power as models reach ceiling performance, highlighting that the remaining useful evaluation relies on scarce expert human…
cs.CY2026
Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts
Alexander K. Saeri, Jess Graham, Michael Noetel +185
Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…