3 papers
cs.LG2025
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
Boyang Zhang, Daning Cheng, Yunquan Zhang +3
The exponential growth in parameter size and computational complexity of deep models poses significant challenges for efficient deployment. The core problem of existing compression…
cs.LG2025
Lossless Model Compression via Joint Low-Rank Factorization Optimization
Boyang Zhang, Daning Cheng, Yunquan Zhang +2
Low-rank factorization is a popular model compression technique that minimizes the error between approximated and original weight matrices. Despite achieving performances clos…
cs.LG2025
FP=xINT:Representing Neural Networks via Low-Bit Series Basis Functions
Boyang Zhang, Daning Cheng, Yunquan Zhang +3
Post-Training Quantization (PTQ) converts pre-trained Full-Precision (FP) models into quantized versions without training. While existing methods reduce size and computational cost…