11 papers
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe +57
Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances,…
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
Dasol Choi, Jihwan Lee, Minjae Lee +1
While prior research on text-to-image generation has predominantly focused on biases in human depictions, demographic bias in generated objects remains relatively underexplored. We…
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
Charvi Rastogi, Mukul Bhutani, Minsuk Kahng +13
Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely calibrated to a Western-centric default, creating significant vulnerabilities for t…
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
Minji Jung, Minjae Lee, Yejin Kim +2
LLM leaderboards are widely used to compare models and guide deployment decisions. However, leaderboard rankings are shaped by evaluation priorities set by benchmark designers, rat…
Data-Prompt Co-Evolution: Growing Test Sets to Refine LLM Behavior
Minjae Lee, Minsuk Kahng
Large Language Models (LLMs) are increasingly embedded in applications, and people can shape model behavior by editing prompt instructions. Yet encoding subtle, domain-specific pol…
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
Lujain Ibrahim, Canfer Akbulut, Rasmi Elasmar +7
The tendency of users to anthropomorphise large language models (LLMs) is of growing interest to AI developers, researchers, and policy-makers. Here, we present a novel method for…