9 papers
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
Nicole Mitchell, Dhruv Agarwal, Maty Bohacek +2
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features…
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe +57
Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances,…
Detecting and Controlling Sycophancy with Cascading Linear Features
Maty Bohacek, Rishub Jain, Nicholas Dufour +3
Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavior. Thes…
Positive Alignment: Artificial Intelligence for Human Flourishing
Ruben Laukkonen, Seb Krier, Chloé Bakalar +13
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psych…
Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity
Pushkar Mishra, Charvi Rastogi, Stephen R. Pfohl +9
Ensuring the safety of Generative AI requires a nuanced understanding of pluralistic viewpoints. In this paper, we introduce a novel data-driven approach for analyzing ordinal safe…
Value Profiles for Encoding Human Variation
Taylor Sorensen, Pushkar Mishra, Roma Patel +6
Modelling human variation in rating tasks is crucial for personalization, pluralistic model alignment, and computational social science. We propose representing individuals using n…