8 papers
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe +57
Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances,…
Land cover and flood type govern the detection limits of satellite-based flood mapping across diverse global flood events
Venkatesh Kolluru, Rajat Shinde, Abdelhak Marouane +6
Floods are among the most destructive natural hazards, and their increasing frequency under climate change makes satellite-based inundation mapping essential for disaster response.…
From NVSS to RACS: Identifying truly Compact and Steep spectrum Radio sources
Rajat Shinde, Yogesh Maan, Apurba Bera
Compact, steep-spectrum radio sources are key tracers of exotic astrophysical objects such as pulsars and high-redshift radio galaxies. All-sky radio surveys at different frequenci…
The CitizenQuery Benchmark: A Novel Dataset and Evaluation Pipeline for Measuring LLM Performance in Citizen Query Tasks
Neil Majithia, Rajat Shinde, Zo Chapman +5
"Citizen queries" are questions asked by an individual about government policies, guidance, and services that are relevant to their circumstances, encompassing a range of topics in…
A refined model of secondary photon emission from heavy WIMP annihilations in the Galactic Centre
Rajat Shinde, Julia Djuvsland, Davide Dapaoli +1
Heavy Weakly Interacting Massive Particles (WIMPs) remain a prominent yet less constrained dark matter (DM) candidate, with the Galactic Centre (GC) serving as a prime target for i…
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…