10 papers
Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment
Arkadiy Saakyan, Charvi Rastogi, Lora Aroyo
Safe global deployment of AI models requires alignment with human values that vary across cultures. Yet rater pools in safety evaluation datasets remain largely geographically homo…
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
Charvi Rastogi, Mukul Bhutani, Minsuk Kahng +13
Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely calibrated to a Western-centric default, creating significant vulnerabilities for t…
Evaluating Language Models for Harmful Manipulation
Canfer Akbulut, Rasmi Elasmar, Abhishek Roy +9
Interest in the concept of AI-driven harmful manipulation is growing, yet current approaches to evaluating it are limited. This paper introduces a framework for evaluating harmful…
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
Lujain Ibrahim, Canfer Akbulut, Rasmi Elasmar +7
The tendency of users to anthropomorphise large language models (LLMs) is of growing interest to AI developers, researchers, and policy-makers. Here, we present a novel method for…
Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity
Pushkar Mishra, Charvi Rastogi, Stephen R. Pfohl +9
Ensuring the safety of Generative AI requires a nuanced understanding of pluralistic viewpoints. In this paper, we introduce a novel data-driven approach for analyzing ordinal safe…
To ArXiv or not to ArXiv: A Study Quantifying Pros and Cons of Posting Preprints Online
Charvi Rastogi, Ivan Stelmakh, Xinwei Shen +4
Double-blind conferences have engaged in debates over whether to allow authors to post their papers online on arXiv or elsewhere during the review process. Independently, some auth…