8 papers
On the Necessity of Output Distribution Reweighting for Effective Class Unlearning
Ali Ebrahimpour-Boroojeny, Yian Wang, Hari Sundaram
In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause information leakage about the forgotten clas…
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
Soorya Ram Shimgekar, Agam Goyal, Amruta Parulekar +6
Large language models (LLMs) are increasingly deployed in conversational settings where user tone ranges from polite to adversarial or toxic, yet less is known about whether toxic…
State Contamination in Memory-Augmented LLM Agents
Yian Wang, Agam Goyal, Yuen Chen +1
LLM agents increasingly rely on persistent state, including transcripts, summaries, retrieved context, and memory buffers, to support long-horizon interaction. This makes safety de…
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
Yian Wang, Yuen Chen, Agam Goyal +1
Large language models (LLMs) frequently generate toxic content, posing significant risks for safe deployment. Current mitigation strategies often degrade generation quality or requ…
From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
Agam Goyal, Yian Wang, Eshwar Chandrasekharan +1
LLM-based social simulations can generate believable community interactions, enabling ``policy wind tunnels'' where governance interventions are tested before deployment. But belie…
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
Agam Goyal, Vedant Rathi, William Yeh +3
Large language models (LLMs) are now ubiquitous in user-facing applications, yet they still generate undesirable toxic outputs, including profanity, vulgarity, and derogatory remar…