3 papers
cs.CL2026
RSRA: Training-Free Probing of Representation Sensitivity for Efficient LoRA Rank Allocation
Jiaqi Liu, Haidong Kang, Qihui Zhao +2
Parameter-efficient fine-tuning enables large language models to adapt to downstream tasks with substantially lower computational and storage cost, and Low-Rank Adaptation (LoRA) i…
cs.LG2025
Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
Haidong Kang, Lianbo Ma, Guo Yu +1
Mixed precision quantization (MPQ) is an effective quantization approach to achieve accuracy-complexity trade-off of neural network, through assigning different bit-widths to netwo…
cs.LG2024
One-Step Forward and Backtrack: Overcoming Zig-Zagging in Loss-Aware Quantization Training
Lianbo Ma, Yuee Zhou, Jianlun Ma +2
Weight quantization is an effective technique to compress deep neural networks for their deployment on edge devices with limited resources. Traditional loss-aware quantization meth…