Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
Weizhong Huang, Yuxin Zhang, Xiawu Zheng +2
In this paper, we address the challenge of determining the layer-wise sparsity rates of large language models (LLMs) through a theoretical perspective. Specifically, we identify a…
cs.LG2025
Towards Efficient Automatic Self-Pruning of Large Language Models
Weizhong Huang, Yuxin Zhang, Xiawu Zheng +2
Despite exceptional capabilities, Large Language Models (LLMs) still face deployment challenges due to their enormous size. Post-training structured pruning is a promising solution…
cs.LG2024
Boosting the Cross-Architecture Generalization of Dataset Distillation through an Empirical Study
Lirui Zhao, Yuxin Zhang, Fei Chao +1
The poor cross-architecture generalization of dataset distillation greatly weakens its practical significance. This paper attempts to mitigate this issue through an empirical study…