5 citations · 9 across the 4 of their papers we have counts for
1 paper · 1 filter
Houwen Peng, Kan Wu, Yixuan Wei +17
In this paper, we explore FP8 low-bit data formats for efficient training of large language models (LLMs). Our key insight is that most variables, such as gradients and optimizer s…