4 citations · 4 across the 3 of their papers we have counts for
4 papers
LLM Robustness Leaderboard v1 --Technical report
Pierre Peigné - Lefebvre, Quentin Feuillade-Montixi, Tom David +1
This technical report accompanies the LLM robustness leaderboard published by PRISM Eval for the Paris AI Action Summit. We introduce PRISM Eval Behavior Elicitation Tool (BET), an…
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
Reva Schwartz, Rumman Chowdhury, Akash Kundu +17
Conventional AI evaluation approaches concentrated within the AI stack exhibit systemic limitations for exploring, navigating and resolving the human and societal factors that play…
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…
Robustness tests for biomedical foundation models should tailor to specifications
R. Patrick Xian, Noah R. Baker, Tom David +5
The rise of biomedical foundation models creates new hurdles in model testing and authorization, given their broad capabilities and susceptibility to complex distribution shifts. W…