4 papers
Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression
Xiaohui Wang, Peng Ye, Chenyu Huang +5
With the rise of the fine-tuned-pretrained paradigm, storing numerous fine-tuned models for multi-tasking creates significant storage overhead. Delta compression alleviates this by…
Decouple and Orthogonalize: A Data-Free Framework for LoRA Merging
Shenghe Zheng, Hongzhi Wang, Chenyu Huang +5
With more open-source models available for diverse tasks, model merging has gained attention by combining models into one, reducing training, storage, and inference costs. Current…
Dynamic Base model Shift for Delta Compression
Chenyu Huang, Peng Ye, Shenghe Zheng +4
Transformer-based models with the pretrain-finetune paradigm bring about significant progress, along with the heavy storage and deployment costs of finetuned models on multiple tas…
Seeing Delta Parameters as JPEG Images: Data-Free Delta Compression with Discrete Cosine Transform
Chenyu Huang, Peng Ye, Xiaohui Wang +5
With transformer-based models and the pretrain-finetune paradigm becoming mainstream, the high storage and deployment costs of individual finetuned models on multiple tasks pose cr…