3 papers
cs.CL2024
MINI-LLM: Memory-Efficient Structured Pruning for Large Language Models
Hongrong Cheng, Miao Zhang, Javen Qinfeng Shi
As Large Language Models (LLMs) grow dramatically in size, there is an increasing trend in compressing and speeding up these models. Previous studies have highlighted the usefulnes…
cs.CV2023
Influence Function Based Second-Order Channel Pruning-Evaluating True Loss Changes For Pruning Is Possible Without Retraining
Hongrong Cheng, Miao Zhang, Javen Qinfeng Shi
A challenge of channel pruning is designing efficient and effective criteria to select channels to prune. A widely used criterion is minimal performance degeneration. To accurately…
cs.LG2023
A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations
Hongrong Cheng, Miao Zhang, Javen Qinfeng Shi
Modern deep neural networks, particularly recent large language models, come with massive model sizes that require significant computational and storage resources. To enable the de…