2 papers
cs.LG2026
Optimizers for Diffusion Models: A Controlled Benchmark
Arman Bolatov, Egor Shulgin, David Li +6
Discrete diffusion models now match autoregressive language models on several benchmarks, while the question of how best to train them has received far less attention: the optimize…
cs.LG2022
Shifted Compression Framework: Generalizations and Improvements
Egor Shulgin, Peter Richtárik
Communication is one of the key bottlenecks in the distributed training of large-scale machine learning models, and lossy compression of exchanged information, such as stochastic g…