Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
Hengjie Cao, Zhendong Huang, Mengyi Chen +15
FP4 training promises substantial memory and compute savings for large language models, but remains fragile because blockwise quantization is dictated by extreme activation magnitu…
cs.LG2024
Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized Models
Yubin Shi, Yixuan Chen, Mingzhi Dong +10
Despite their prevalence in deep-learning communities, over-parameterized models convey high demands of computational costs for proper training. This work studies the fine-grained,…