18 citations · 18 across the 2 of their papers we have counts for
1 paper · 1 filter
Tao Yu, Gaurav Gupta, Karthick Gopalswamy +7
Large models training is plagued by the intense compute cost and limited hardware memory. A practical solution is low-precision representation but is troubled by loss in numerical…