2 papers
cs.LG2026
Big2Small: A Unifying Neural Network Framework for Model Compression
Jing-Xiao Liao, Haoran Wang, Tao Li +4
With the development of foundational models, model compression has become a critical requirement. Various model compression approaches have been proposed such as low-rank decomposi…
cs.LG2024
DeltaDQ: Ultra-High Delta Compression for Fine-Tuned LLMs via Group-wise Dropout and Separate Quantization
Yanfeng Jiang, Zelan Yang, Bohua Chen +3
Large language models achieve exceptional performance on various downstream tasks through supervised fine-tuning. However, the diversity of downstream tasks and practical requireme…