4 papers
A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders
Chenhao Zhang, Chris Lin, Su-In Lee
We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs). While SAEs improve interpreta…
Unlearning Evaluation through Subset Statistical Independence
Chenhao Zhang, Muxing Li, Feng Liu +2
Evaluating machine unlearning remains challenging, as existing methods typically require retraining reference models or performing membership inference attacks, both of which rely…
Machine Unlearning for Streaming Forgetting
Shaofei Shen, Chenhao Zhang, Yawen Zhao +3
Machine unlearning aims to remove knowledge of the specific training data in a well-trained model. Currently, machine unlearning methods typically handle all forgetting data in a s…
Toward Efficient Data-Free Unlearning
Chenhao Zhang, Shaofei Shen, Weitong Chen +1
Machine unlearning without access to real data distribution is challenging. The existing method based on data-free distillation achieved unlearning by filtering out synthetic sampl…