1 paper
Alireza Olama, Andreas Lundell, Izzat El Hajj +2
Inter-node communication bandwidth increasingly constrains distributed training at scale on multi-node GPU clusters. While compact models are the ultimate deployment target, conven…