26 citations · 32 across the 7 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
CurvaDion: Curvature-Adaptive Distributed Orthonormalization
Bhavesh Kumar, Roger Jin, Jeffrey Quesnelle
As language models scale to trillions of parameters, distributed training across many GPUs becomes essential, yet gradient synchronization over high-bandwidth, low-latency networks…
cs.LG2024
DeMo: Decoupled Momentum Optimization
Bowen Peng, Lizhang Chen, Baiyu Su +3
Scaling neural network training increasingly depends on synchronous data-parallelism, yet full-precision gradient all-reduce imposes a severe communication bottleneck. We propose D…