3 papers
cs.CL2026
GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation
Hao Liu, Guangyan Li, Wensheng Zhang +1
Large Language Models (LLMs) exhibit strong reasoning abilities, but their high computational costs limit their practical deployment. Recent studies reveal significant redundancy i…
cs.LG2025
MGAA: Multi-Granular Adaptive Allocation fof Low-Rank Compression of LLMs
Guangyan Li, Yongqiang Tang, Wensheng Zhang
The enormous parameter scale of large language models (LLMs) has made model compression a research hotspot, which aims to alleviate computational resource demands during deployment…
cs.LG2024
LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models
Guangyan Li, Yongqiang Tang, Wensheng Zhang
Large language models (LLMs) show excellent performance in difficult tasks, but they often require massive memories and computational resources. How to reduce the parameter scale o…