6 papers
Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension
Jinhao Zhang, Zeyu Liu, Zicheng Yan +4
Existing theories of neural-network width characterize asymptotic limits, but provide limited guidance on whether an expansion direction identified from finite training data remain…
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
Boyang Zhang, Daning Cheng, Yunquan Zhang +3
The exponential growth in parameter size and computational complexity of deep models poses significant challenges for efficient deployment. The core problem of existing compression…
Compression for Better: A General and Stable Lossless Compression Framework
Boyang Zhang, Daning Cheng, Yunquan Zhang +2
This work focus on how to stabilize and lossless model compression, aiming to reduce model complexity and enhance efficiency without sacrificing performance due to compression erro…
Lossless Model Compression via Joint Low-Rank Factorization Optimization
Boyang Zhang, Daning Cheng, Yunquan Zhang +2
Low-rank factorization is a popular model compression technique that minimizes the error between approximated and original weight matrices. Despite achieving performances clos…
FP=xINT:Representing Neural Networks via Low-Bit Series Basis Functions
Boyang Zhang, Daning Cheng, Yunquan Zhang +3
Post-Training Quantization (PTQ) converts pre-trained Full-Precision (FP) models into quantized versions without training. While existing methods reduce size and computational cost…
Can the capability of Large Language Models be described by human ability? A Meta Study
Mingrui Zan, Yunquan Zhang, Boyang Zhang +2
Users of Large Language Models (LLMs) often perceive these models as intelligent entities with human-like capabilities. However, the extent to which LLMs' capabilities truly approx…