4 papers
Measuring, Localizing, and Ablating Alignment Signatures in LLMs
Aniket Anand, Janvijay Singh, Zhewei Sun +2
Aligned language models often exhibit a recognizable AI-like style, yet its connection to post-training and internal representations remains poorly understood. In this work, we stu…
Governance of AI-Generated Content: A Case Study on Social Media Platforms
Lan Gao, Abani Ahmed, Oscar Chen +6
Online platforms are seeing increasing amounts of AI-generated content -- text and other forms of media that are made or co-created with generative AI. This trend suggests platform…
Beyond PII: How Users Attempt to Estimate and Mitigate Implicit LLM Inference
Synthia Wang, Sai Teja Peddinti, Nina Taft +1
Large Language Models (LLMs) such as ChatGPT can infer personal attributes from seemingly innocuous text, raising privacy risks beyond memorized data leakage. While prior work has…
Can LLMs Address Mental Health Questions? A Comparison with Human Therapists
Synthia Wang, Yuwei Cheng, Austin Song +3
Limited access to mental health care has motivated the use of digital tools and conversational agents powered by large language models (LLMs), yet their quality and reception remai…