3 papers
cs.CY2025
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…
cs.CL2025
Analysis of Indic Language Capabilities in LLMs
Aatman Vaidya, Tarunima Prabhakar, Denny George +1
This report evaluates the performance of text-in text-out Large Language Models (LLMs) to understand and generate Indic languages. This evaluation is used to identify and prioritiz…
cs.CL2024
The Uli Dataset: An Exercise in Experience Led Annotation of oGBV
Arnav Arora, Maha Jinadoss, Cheshta Arora +22
Online gender based violence has grown concomitantly with adoption of the internet and social media. Its effects are worse in the Global majority where many users use social media…