8 papers
On the Necessity of Output Distribution Reweighting for Effective Class Unlearning
Ali Ebrahimpour-Boroojeny, Yian Wang, Hari Sundaram
In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause information leakage about the forgotten clas…
State Contamination in Memory-Augmented LLM Agents
Yian Wang, Agam Goyal, Yuen Chen +1
LLM agents increasingly rely on persistent state, including transcripts, summaries, retrieved context, and memory buffers, to support long-horizon interaction. This makes safety de…
AMUN: Adversarial Machine UNlearning
Ali Ebrahimpour-Boroojeny, Hari Sundaram, Varun Chandrasekaran
Machine unlearning, where users can request the deletion of a forget dataset, is becoming increasingly important because of numerous privacy regulations. Initial works on ``exact''…
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
Agam Goyal, Vedant Rathi, William Yeh +3
Large language models (LLMs) are now ubiquitous in user-facing applications, yet they still generate undesirable toxic outputs, including profanity, vulgarity, and derogatory remar…
DIVEBATCH: Accelerating Model Training Through Gradient-Diversity Aware Batch Size Adaptation
Yuen Chen, Yian Wang, Hari Sundaram
The goal of this paper is to accelerate the training of machine learning models, a critical challenge since the training of large-scale deep neural models can be computationally ex…
The Chilling: Identifying Strategic Antisocial Behavior Online and Examining the Impact on Journalists
Yian Wang, Mukhilshankar Umashankar, Eshwar Chandrasekharan +1
On social platforms like Twitter, strategic targeted attacks are becoming increasingly common, especially against vulnerable groups such as female journalists. Two key challenges i…