1 paper
Jari Kolehmainen, Nikolay Blagoev, Semih Kara +3
Training large language models is generally done on clusters containing thousands of accelerators, communicating over a high-bandwidth interconnect. Scaling up these clusters is ex…