From the 1 of 2 linked papers with an AI index.
2 papers
cs.CY2026
The Benchmark Ceiling: Human Judgment, Evaluation Scarcity, and the Political Economy of AI Capability Measurement
Mark Esposito, Liu Zhang
The paper examines how AI benchmarks lose discriminative power as models reach ceiling performance, highlighting that the remaining useful evaluation relies on scarce expert human…
econ.GN2026
No Last Mile: A Theory of the Human Data Market
Ali Ansari, Mark Esposito, Ava Fitoussy +1
The standard framing treats structured human-data work as transitional, a bridge between today's imperfect models and a future state where automation is complete. We challenge this…