2 papers
cs.CV2024
TT-MPD: Test Time Model Pruning and Distillation
Haihang Wu, Wei Wang, Tamasha Malepathirana +3
Pruning can be an effective method of compressing large pre-trained models for inference speed acceleration. Previous pruning approaches rely on access to the original training dat…
cs.CL2024
LLM-BIP: Structured Pruning for Large Language Models with Block-Wise Forward Importance Propagation
Haihang Wu
Large language models (LLMs) have demonstrated remarkable performance across various language tasks, but their widespread deployment is impeded by their large size and high computa…