5 citations · 5 across the 12 of their papers we have counts for
1 paper · 1 filter
Johannes Hagemann, Samuel Weinbach, Konstantin Dobler +2
Efficiently training large language models requires parallelizing across hundreds of hardware accelerators and invoking various compute and memory optimizations. When combined, man…