Mosaic: Composite Projection Pruning for Resource-efficient LLMs
arXiv:2504.06323 · doi:10.1016/j.future.2025.108056
Abstract
Extensive compute and memory requirements limit the deployment of large language models (LLMs) on any hardware. Compression methods, such as pruning, can reduce model size, which in turn reduces resource requirements. State-of-the-art pruning is based on coarse-grained methods. They are time-consuming and inherently remove critical model parameters, adversely impacting the quality of the pruned model. This paper introduces projection pruning, a novel fine-grained method for pruning LLMs. In addition, LLM projection pruning is enhanced by a new approach we refer to as composite projection pruning - the synergistic combination of unstructured pruning that retains accuracy and structured pruning that reduces model size. We develop Mosaic, a novel system to create and deploy pruned LLMs using composite projection pruning. Mosaic is evaluated using a range of performance and quality metrics on multiple hardware platforms, LLMs, and datasets. Mosaic is 7.19x faster in producing models than existing approaches. Mosaic models achieve up to 84.2% lower perplexity and 31.4% higher accuracy than models obtained from coarse-grained pruning. Up to 67% faster inference and 68% lower GPU memory use is noted for Mosaic models. Mosaic is available for public use from https://github.com/blessonvar/Mosaic
References in corpus (13)
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Gemini: A Family of Highly Capable Multimodal Models
- Gemma: Open Models Based on Gemini Research and Technology
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities
- DNNShifter: An Efficient DNN Pruning System for Edge Computing
- Structured Pruning is All You Need for Pruning CNNs at Initialization
- MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
- Compact Language Models via Pruning and Knowledge Distillation
- Nemotron-4 340B Technical Report
- Rapid Deployment of DNNs for Edge Computing via Structured Pruning at Initialization
- RakutenAI-7B: Extending Large Language Models for Japanese