9 papers
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
Hengjie Cao, Zhendong Huang, Mengyi Chen +15
FP4 training promises substantial memory and compute savings for large language models, but remains fragile because blockwise quantization is dictated by extreme activation magnitu…
Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics
Hai Wang, Xiaochen Yang, Mingzhi Dong +1
The dream of instantly creating rich 360-degree panoramic worlds from text is rapidly becoming a reality, yet a crucial gap exists in our ability to reliably evaluate their semanti…
Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers
Anrui Chen, Ruijun Huang, Xin Zhang +15
Mixture-of-Experts (MoE) architectures are often considered a natural fit for continual learning because sparse routing should localize updates and reduce interference, yet MoE Tra…
SD-MoE: Spectral Decomposition for Effective Expert Specialization
Ruijun Huang, Fang Dong, Xin Zhang +16
Mixture-of-Experts (MoE) architectures scale Large Language Models via expert specialization induced by conditional computation. In practice, however, expert specialization often f…
Dispelling the Curse of Singularities in Neural Network Optimizations
Hengjie Cao, Mengyi Chen, Yifeng Yang +11
This work investigates the optimization instability of deep neural networks from a less-explored yet insightful perspective: the emergence and amplification of singularities in the…
Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy
Zhendong Huang, Hengjie Cao, Fang Dong +14
Gradient signals in LLM training are highly anisotropic: recurrent linguistic structure concentrates energy into a small set of dominant spectral directions, while context specific…