2 citations · 5 across the 8 of their papers we have counts for
4 papers · 1 filter
What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Meera Desai, Sang T. Truong, Hanna Wallach +8
Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reason…
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
Cristian Trout, Sanmi Koyejo, Sasha Romanosky +34
From maritime trade to commercial nuclear power, insurance has been the enabler of major economic and technological developments by pricing risk, limiting downside, and spreading b…
Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming
Alexandra Chouldechova, A. Feder Cooper, Solon Barocas +3
We argue that conclusions drawn about relative system safety or attack method efficacy via AI red teaming are often not supported by evidence provided by attack success rate (ASR)…
Extracting books from production language models
Ahmed Ahmed, A. Feder Cooper, Sanmi Koyejo +1
Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model's weights during training, and whether…