21 papers
Dissociating the Internal Representations of Sycophancy in LLMs
Anthony Baez, Sheer Karny, Pat Pataranutaporn
Large Language Models (LLMs) frequently exhibit sycophancy, agreeing with a user's statement even when it is incorrect. While often studied as a single, uniform behavior, sycophanc…
Multi-Turn Neural Transparency: Surfacing Neural Activations Improves User Calibration to LLM Behavioral Drift
Sheer Karny, Anthony Baez, Pat Pataranutaporn
Chatbot behavior is often opaque to users, as responses can shift unpredictably across a conversation, drifting toward sycophancy, toxicity, or other unsafe responses. This can lea…
AI-Wrapped: Participatory, Privacy-Preserving Measurement of Longitudinal LLM Use In-the-Wild
Cathy Mengying Fang, Sheer Karny, Chayapatr Archiwaranguprok +3
Alignment research on large language models (LLMs) increasingly depends on understanding how these systems are used in everyday contexts. Yet naturalistic interaction data is diffi…
AI-Generated Letters from the Future: A Randomized Test of Personalized Climate Communication
Nattavudh Powdthavee, Pat Pataranutaporn, Sandra J. Geiger +2
We examined whether personalized, AI-generated letters from the future can increase public engagement with climate action. In a preregistered online experiment with 1,654 U.S. pare…
"Death" of a Chatbot: Investigating and Designing Toward Psychologically Safe Endings for Human-AI Relationships
Rachel Poonsiriwong, Chayapatr Archiwaranguprok, Pat Pataranutaporn
Millions of users form emotional attachments to AI companions like Character AI, Replika, and ChatGPT. When these relationships end through model updates, safety interventions, or…
Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures
Hua Shen, Tiffany Knearem, Divy Thakkar +9
The rapid integration of generative AI into everyday life underscores the need to move beyond unidirectional alignment models that only adapt AI to human values. This workshop focu…