2 citations · 8 across the 11 of their papers we have counts for
Showing 2026 · cs.CYShow all
2 papers · 2 filters
cs.CY2026
What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Meera Desai, Sang T. Truong, Hanna Wallach +8
Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reason…
cs.CY2026
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
Cristian Trout, Sanmi Koyejo, Sasha Romanosky +34
From maritime trade to commercial nuclear power, insurance has been the enabler of major economic and technological developments by pricing risk, limiting downside, and spreading b…