3 papers
cs.LG2026
Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension
Jinhao Zhang, Zeyu Liu, Zicheng Yan +4
Existing theories of neural-network width characterize asymptotic limits, but provide limited guidance on whether an expansion direction identified from finite training data remain…
cs.LG2026
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
Jinhao Zhang, Yunquan Zhang, Zicheng yan +3
Post Training Quantization (PTQ), a mainstream model compression technique, often leads to the paradoxical 'low error, high loss' phenomenon because it focuses solely on minimizing…
cs.LG2026
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
Jinhao Zhang, Yunquan Zhang, Daning Chen +2
Current mainstream post-training quantization methods for large language models typically apply a uniform quantization strategy across all network layers, overlooking the substanti…