5 citations · 7 across the 17 of their papers we have counts for
6 papers · 1 filter
Towards the Connection between Activation Sparsity and Flat Minima
Ze Peng, Jian Zhang, Lei Qi +2
The observation that activation sparsity emerges in MLP blocks of standardly trained Transformers offers an opportunity to drastically reduce computation costs without sacrificing…
Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGD
Ze Peng, Jian Zhang, Yisen Wang +3
Information-theoretic (IT) generalization bounds have been used to study the generalization of learning algorithms. These bounds are intrinsically data- and algorithm-dependent so…
MAGIC: Achieving Superior Model Merging via Magnitude Calibration
Yayuan Li, Jian Zhang, Jintao Guo +4
The proliferation of pre-trained models has given rise to a wide array of specialised, fine-tuned models. Model merging aims to merge the distinct capabilities of these specialised…
On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning
Ze Peng, Jian Zhang, Jintao Guo +3
Continual learning seeks the human-like ability to accumulate new skills in machine intelligence. Its central challenge is catastrophic forgetting, whose underlying cause has not b…
PG-LBO: Enhancing High-Dimensional Bayesian Optimization with Pseudo-Label and Gaussian Process Guidance
Taicai Chen, Yue Duan, Dong Li +3
Variational Autoencoder based Bayesian Optimization (VAE-BO) has demonstrated its excellent performance in addressing high-dimensional structured optimization problems. However, cu…
A Theoretical Explanation of Activation Sparsity through Flat Minima and Adversarial Robustness
Ze Peng, Lei Qi, Yinghuan Shi +1
A recent empirical observation (Li et al., 2022b) of activation sparsity in MLP blocks offers an opportunity to drastically reduce computation costs for free. Although having attri…