44 citations · 44 across the 1 of their papers we have counts for
1 paper · 1 filter
Gyudong Kim, Hyukju Na, Jin Hyeon Kim +6
As training billion-scale transformers becomes increasingly common, employing multiple distributed GPUs along with parallel training methods has become a standard practice. However…