From the 1 of 3 linked papers with an AI index.
3 papers
cs.CY2026
The Benchmark Ceiling: Human Judgment, Evaluation Scarcity, and the Political Economy of AI Capability Measurement
Mark Esposito, Liu Zhang
The paper examines how AI benchmarks lose discriminative power as models reach ceiling performance, highlighting that the remaining useful evaluation relies on scarce expert human…
cs.CY2026
Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts
Alexander K. Saeri, Jess Graham, Michael Noetel +185
Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…
econ.GN2026
No Last Mile: A Theory of the Human Data Market
Ali Ansari, Mark Esposito, Ava Fitoussy +1
The standard framing treats structured human-data work as transitional, a bridge between today's imperfect models and a future state where automation is complete. We challenge this…