2 papers
cs.CY2025
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…
cs.CL2025
MSTS: A Multimodal Safety Test Suite for Vision-Language Models
Paul Röttger, Giuseppe Attanasio, Felix Friedrich +19
Vision-language models (VLMs), which process image and text inputs, are increasingly integrated into chat assistants and other consumer AI applications. Without proper safeguards,…