From the 1 of 5 linked papers with an AI index.
5 papers
Deep Simulation-Based Inference for Inhomogeneous Bivariate Log-Gaussian Cox Processes
Qihan Zou, Yan Wang, Tingjin Chu +1
The paper presents a two‑step simulation‑based estimation approach that first fits Poisson first‑order parameters and then uses neural networks to infer latent field parameters of…
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
Hengjie Cao, Zhendong Huang, Mengyi Chen +15
FP4 training promises substantial memory and compute savings for large language models, but remains fragile because blockwise quantization is dictated by extreme activation magnitu…
Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers
Anrui Chen, Ruijun Huang, Xin Zhang +15
Mixture-of-Experts (MoE) architectures are often considered a natural fit for continual learning because sparse routing should localize updates and reduce interference, yet MoE Tra…
SD-MoE: Spectral Decomposition for Effective Expert Specialization
Ruijun Huang, Fang Dong, Xin Zhang +16
Mixture-of-Experts (MoE) architectures scale Large Language Models via expert specialization induced by conditional computation. In practice, however, expert specialization often f…
Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy
Zhendong Huang, Hengjie Cao, Fang Dong +14
Gradient signals in LLM training are highly anisotropic: recurrent linguistic structure concentrates energy into a small set of dominant spectral directions, while context specific…