17 citations · 60 across the 22 of their papers we have counts for
30 papers
AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations
Rachel Poonsiriwong, Chayapatr, Archiwaranguprok +4
Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI Watchdog, a browser-based ag…
Dissociating the Internal Representations of Sycophancy in LLMs
Anthony Baez, Sheer Karny, Pat Pataranutaporn
Large Language Models (LLMs) frequently exhibit sycophancy, agreeing with a user's statement even when it is incorrect. While often studied as a single, uniform behavior, sycophanc…
Multi-Turn Neural Transparency: Surfacing Neural Activations Improves User Calibration to LLM Behavioral Drift
Sheer Karny, Anthony Baez, Pat Pataranutaporn
Chatbot behavior is often opaque to users, as responses can shift unpredictably across a conversation, drifting toward sycophancy, toxicity, or other unsafe responses. This can lea…
AI-Wrapped: Participatory, Privacy-Preserving Measurement of Longitudinal LLM Use In-the-Wild
Cathy Mengying Fang, Sheer Karny, Chayapatr Archiwaranguprok +3
Alignment research on large language models (LLMs) increasingly depends on understanding how these systems are used in everyday contexts. Yet naturalistic interaction data is diffi…
AI-Generated Letters from the Future: A Randomized Test of Personalized Climate Communication
Nattavudh Powdthavee, Pat Pataranutaporn, Sandra J. Geiger +2
We examined whether personalized, AI-generated letters from the future can increase public engagement with climate action. In a preregistered online experiment with 1,654 U.S. pare…
"Death" of a Chatbot: Investigating and Designing Toward Psychologically Safe Endings for Human-AI Relationships
Rachel Poonsiriwong, Chayapatr Archiwaranguprok, Pat Pataranutaporn
Millions of users form emotional attachments to AI companions like Character AI, Replika, and ChatGPT. When these relationships end through model updates, safety interventions, or…