activity
20242026
collaborators

10 papers

cs.CY2026

Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment

Arkadiy Saakyan, Charvi Rastogi, Lora Aroyo

Safe global deployment of AI models requires alignment with human values that vary across cultures. Yet rater pools in safety evaluation datasets remain largely geographically homo…

cs.CY2026

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South

Charvi Rastogi, Mukul Bhutani, Minsuk Kahng +13

Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely calibrated to a Western-centric default, creating significant vulnerabilities for t…

cs.AI2026

Evaluating Language Models for Harmful Manipulation

Canfer Akbulut, Rasmi Elasmar, Abhishek Roy +9

Interest in the concept of AI-driven harmful manipulation is growing, yet current approaches to evaluating it are limited. This paper introduces a framework for evaluating harmful…

cs.CL2026

Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models

Lujain Ibrahim, Canfer Akbulut, Rasmi Elasmar +7

The tendency of users to anthropomorphise large language models (LLMs) is of growing interest to AI developers, researchers, and policy-makers. Here, we present a novel method for…

cs.CY2026

Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity

Pushkar Mishra, Charvi Rastogi, Stephen R. Pfohl +9

Ensuring the safety of Generative AI requires a nuanced understanding of pluralistic viewpoints. In this paper, we introduce a novel data-driven approach for analyzing ordinal safe…

cs.DL2025

To ArXiv or not to ArXiv: A Study Quantifying Pros and Cons of Posting Preprints Online

Charvi Rastogi, Ivan Stelmakh, Xinwei Shen +4

Double-blind conferences have engaged in debates over whether to allow authors to post their papers online on arXiv or elsewhere during the review process. Independently, some auth…