4 papers
Towards the Connection between Activation Sparsity and Flat Minima
Ze Peng, Jian Zhang, Lei Qi +2
The observation that activation sparsity emerges in MLP blocks of standardly trained Transformers offers an opportunity to drastically reduce computation costs without sacrificing…
When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging
Yayuan Li, Ze Peng, Jian Zhang +3
Model merging combines multiple fine-tuned models into a single model by adding their weight updates, providing a lightweight alternative to retraining. Existing methods primarily…
Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGD
Ze Peng, Jian Zhang, Yisen Wang +3
Information-theoretic (IT) generalization bounds have been used to study the generalization of learning algorithms. These bounds are intrinsically data- and algorithm-dependent so…
On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning
Ze Peng, Jian Zhang, Jintao Guo +3
Continual learning seeks the human-like ability to accumulate new skills in machine intelligence. Its central challenge is catastrophic forgetting, whose underlying cause has not b…