4 papers
Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
Yihong Huang, Fei Ma, Yihua Shao +4
Vision token pruning has proven to be an effective acceleration technique for the efficient Vision Language Model (VLM). However, existing pruning methods demonstrate excellent per…
ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation
Yihua Shao, Xiaofeng Lin, Xinwei Long +7
Enabling multi-task adaptation in pre-trained Low-Rank Adaptation (LoRA) models is crucial for enhancing their generalization capabilities. Most existing pre-trained LoRA fusion me…
In-Context Meta LoRA Generation
Yihua Shao, Minxi Yan, Yang Liu +12
Low-rank Adaptation (LoRA) has demonstrated remarkable capabilities for task specific fine-tuning. However, in scenarios that involve multiple tasks, training a separate LoRA model…
GWQ: Gradient-Aware Weight Quantization for Large Language Models
Yihua Shao, Yan Gu, Siyu Chen +12
Large language models (LLMs) show impressive performance in solving complex language tasks. However, its large number of parameters presents significant challenges for the deployme…