6 papers
Quantifying Depth Sufficiency in Residual Neural Networks: A First-Order Criterion
Zeyu Liu, Jinhao Zhang, Yunquan Zhang +4
How can we determine whether a trained neural network is already deep enough? We study this under a fixed function-preserving residual-growth protocol specifying insertion location…
Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension
Jinhao Zhang, Zeyu Liu, Zicheng Yan +4
Existing theories of neural-network width characterize asymptotic limits, but provide limited guidance on whether an expansion direction identified from finite training data remain…
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
Jinhao Zhang, Yunquan Zhang, Zicheng yan +3
Post Training Quantization (PTQ), a mainstream model compression technique, often leads to the paradoxical 'low error, high loss' phenomenon because it focuses solely on minimizing…
Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix
Jinhao Zhang, Kangfei Zhao, Qiuhao Zeng +1
Transformer-based architectures have become the dominant paradigm for Continuous-Time Dynamic Graph (CTDG) learning, yet their performance remains limited on temporally shifted dat…
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
Jinhao Zhang, Yunquan Zhang, Daning Chen +2
Current mainstream post-training quantization methods for large language models typically apply a uniform quantization strategy across all network layers, overlooking the substanti…
MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
Jinhao Zhang, Yunquan Zhang, Boyang Zhang +2
Quantization method plays a crucial role in improving model efficiency and reducing deployment costs, enabling the widespread application of deep learning models on resource-constr…