activity
20242026
collaborators

21 papers

cs.LG2026

Dissociating the Internal Representations of Sycophancy in LLMs

Anthony Baez, Sheer Karny, Pat Pataranutaporn

Large Language Models (LLMs) frequently exhibit sycophancy, agreeing with a user's statement even when it is incorrect. While often studied as a single, uniform behavior, sycophanc…

cs.HC2026

Multi-Turn Neural Transparency: Surfacing Neural Activations Improves User Calibration to LLM Behavioral Drift

Sheer Karny, Anthony Baez, Pat Pataranutaporn

Chatbot behavior is often opaque to users, as responses can shift unpredictably across a conversation, drifting toward sycophancy, toxicity, or other unsafe responses. This can lea…

cs.HC2026

AI-Wrapped: Participatory, Privacy-Preserving Measurement of Longitudinal LLM Use In-the-Wild

Cathy Mengying Fang, Sheer Karny, Chayapatr Archiwaranguprok +3

Alignment research on large language models (LLMs) increasingly depends on understanding how these systems are used in everyday contexts. Yet naturalistic interaction data is diffi…

cs.CY2026

AI-Generated Letters from the Future: A Randomized Test of Personalized Climate Communication

Nattavudh Powdthavee, Pat Pataranutaporn, Sandra J. Geiger +2

We examined whether personalized, AI-generated letters from the future can increase public engagement with climate action. In a preregistered online experiment with 1,654 U.S. pare…

cs.HC2026

"Death" of a Chatbot: Investigating and Designing Toward Psychologically Safe Endings for Human-AI Relationships

Rachel Poonsiriwong, Chayapatr Archiwaranguprok, Pat Pataranutaporn

Millions of users form emotional attachments to AI companions like Character AI, Replika, and ChatGPT. When these relationships end through model updates, safety interventions, or…

cs.HC2025

Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures

Hua Shen, Tiffany Knearem, Divy Thakkar +9

The rapid integration of generative AI into everyday life underscores the need to move beyond unidirectional alignment models that only adapt AI to human values. This workshop focu…