3 papers
cs.CY2025
Ask What Your Country Can Do For You: Towards a Public Red Teaming Model
Wm. Matthew Kennedy, Cigdem Patlak, Jayraj Dave +10
AI systems have the potential to produce both benefits and harms, but without rigorous and ongoing adversarial evaluation, AI actors will struggle to assess the breadth and magnitu…
cs.IR2025
Cascade! Human in the loop shortcomings can increase the risk of failures in recommender systems
Wm. Matthew Kennedy, Nishanshi Shukla, Cigdem Patlak +5
Recommender systems are among the most commonly deployed systems today. Systems design approaches to AI-powered recommender systems have done well to urge recommender system develo…
cs.CY2025
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…