14 papers
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
Jinhao Zhang, Yunquan Zhang, Zicheng yan +3
Post Training Quantization (PTQ), a mainstream model compression technique, often leads to the paradoxical 'low error, high loss' phenomenon because it focuses solely on minimizing…
Proximal-Based Generative Modeling for Bayesian Inverse Problems
Boyang Zhang, Zhiguo Wang, Ya-Feng Liu
Score-based diffusion models demonstrate superior performance in generative tasks but encounter fundamental bottlenecks in inverse problems due to the analytical intractability of…
A Qualitative Test-Risk Mechanism for Scaling Behavior in Normalized Residual Networks
Daning Cheng, Zeyu Liu, Jun Sun +4
The scaling behavior, in which test performance often improves as model size and data increase, is a central empirical phenomenon in modern deep learning, yet its theoretical basis…
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
Xin Ye, Daning Cheng, Boyang Zhang +1
Training large-scale Mixture-of-Experts (MoE) models typically requires high-memory, high-bandwidth GPUs (e.g., A100), and their high cost has become a major barrier to large-model…
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
Boyang Zhang, Daning Cheng, Yunquan Zhang +3
The exponential growth in parameter size and computational complexity of deep models poses significant challenges for efficient deployment. The core problem of existing compression…
Compression for Better: A General and Stable Lossless Compression Framework
Boyang Zhang, Daning Cheng, Yunquan Zhang +2
This work focus on how to stabilize and lossless model compression, aiming to reduce model complexity and enhance efficiency without sacrificing performance due to compression erro…