collaborators

8 papers

cs.LG2026

On the Necessity of Output Distribution Reweighting for Effective Class Unlearning

Ali Ebrahimpour-Boroojeny, Yian Wang, Hari Sundaram

In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause information leakage about the forgotten clas…

cs.CL2026

Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits

Soorya Ram Shimgekar, Agam Goyal, Amruta Parulekar +6

Large language models (LLMs) are increasingly deployed in conversational settings where user tone ranges from polite to adversarial or toxic, yet less is known about whether toxic…

cs.AI2026

State Contamination in Memory-Augmented LLM Agents

Yian Wang, Agam Goyal, Yuen Chen +1

LLM agents increasingly rely on persistent state, including transcripts, summaries, retrieved context, and memory buffers, to support long-horizon interaction. This makes safety de…

cs.CL2026

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification

Yian Wang, Yuen Chen, Agam Goyal +1

Large language models (LLMs) frequently generate toxic content, posing significant risks for safe deployment. Current mitigation strategies often degrade generation quality or requ…

cs.CL2026

From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities

Agam Goyal, Yian Wang, Eshwar Chandrasekharan +1

LLM-based social simulations can generate believable community interactions, enabling ``policy wind tunnels'' where governance interventions are tested before deployment. But belie…

cs.CL2025

Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders

Agam Goyal, Vedant Rathi, William Yeh +3

Large language models (LLMs) are now ubiquitous in user-facing applications, yet they still generate undesirable toxic outputs, including profanity, vulgarity, and derogatory remar…