1 paper
Saaketh Narayan, Abhay Gupta, Mansheej Paul +1
Large Language Model training with 8-bit floating point (FP8) formats promises significant efficiency improvements, but reduced numerical precision makes training challenging. It i…