3 papers
cs.LG2025
AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent
Jing Liu, Toshiaki Koike-Akino, Ye Wang +2
To address the enormous size of Large Language Models (LLMs), model compression methods, such as quantization and pruning, are often deployed, especially on edge devices. In this w…
cs.LG2025
TuneComp: Joint Fine-tuning and Compression for Large Foundation Models
Xiangyu Chen, Jing Liu, Ye Wang +4
To reduce model size during post-training, compression methods, including knowledge distillation, low-rank approximation, and pruning, are often applied after fine-tuning the model…
cs.LG2025
LatentLLM: Attention-Aware Joint Tensor Compression
Toshiaki Koike-Akino, Xiangyu Chen, Jing Liu +4
Modern foundation models such as large language models (LLMs) and large multi-modal models (LMMs) require a massive amount of computational and memory resources. We propose a new f…