collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

Towards the Connection between Activation Sparsity and Flat Minima

Ze Peng, Jian Zhang, Lei Qi +2

The observation that activation sparsity emerges in MLP blocks of standardly trained Transformers offers an opportunity to drastically reduce computation costs without sacrificing…

cs.LG2026

When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging

Yayuan Li, Ze Peng, Jian Zhang +3

Model merging combines multiple fine-tuned models into a single model by adding their weight updates, providing a lightweight alternative to retraining. Existing methods primarily…

cs.LG2026

Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGD

Ze Peng, Jian Zhang, Yisen Wang +3

Information-theoretic (IT) generalization bounds have been used to study the generalization of learning algorithms. These bounds are intrinsically data- and algorithm-dependent so…

cs.LG2025

On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning

Ze Peng, Jian Zhang, Jintao Guo +3

Continual learning seeks the human-like ability to accumulate new skills in machine intelligence. Its central challenge is catastrophic forgetting, whose underlying cause has not b…

cs.LG2023

A Theoretical Explanation of Activation Sparsity through Flat Minima and Adversarial Robustness

Ze Peng, Lei Qi, Yinghuan Shi +1

A recent empirical observation (Li et al., 2022b) of activation sparsity in MLP blocks offers an opportunity to drastically reduce computation costs for free. Although having attri…