Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
Jiecheng Zhou, Ding Tang, Rong Fu +8
The burgeoning computational demands for training large language models (LLMs) necessitate efficient methods, including quantized training, which leverages low-bit arithmetic opera…
cs.LG2024
PackMamba: Efficient Processing of Variable-Length Sequences in Mamba training
Haoran Xu, Ziqian Liu, Rong Fu +5
With the evolution of large language models, traditional Transformer models become computationally demanding for lengthy sequences due to the quadratic growth in computation with r…