4 papers · 1 filter
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
Nicole Mitchell, Dhruv Agarwal, Maty Bohacek +2
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features…
Detecting and Controlling Sycophancy with Cascading Linear Features
Maty Bohacek, Rishub Jain, Nicholas Dufour +3
Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavior. Thes…
Positive Alignment: Artificial Intelligence for Human Flourishing
Ruben Laukkonen, Seb Krier, Chloé Bakalar +13
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psych…
Insights on Disagreement Patterns in Multimodal Safety Perception across Diverse Rater Groups
Charvi Rastogi, Tian Huey Teh, Pushkar Mishra +10
AI systems crucially rely on human ratings, but these ratings are often aggregated, obscuring the inherent diversity of perspectives in real-world phenomenon. This is particularly…