Showing 2024Show all
2 papers · 1 filter
cs.AI2024
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
Amelia Hardy, Anka Reuel, Kiana Jafari Meimandi +6
Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market. However, these advertised scores do not…
cs.CY2024
GPAI Evaluations Standards Taskforce: Towards Effective AI Governance
Patricia Paskov, Lukas Berglund, Everett Smith +1
General-purpose AI evaluations have been proposed as a promising way of identifying and mitigating systemic risks posed by AI development and deployment. While GPAI evaluations pla…