Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Scalable Hyperparameter-Divergent Ensemble Training with Automatic Learning Rate Exploration for Large Models
Hailing Cheng, Tao Huang, Chen Zhu +1
Training large neural networks with data-parallel stochastic gradient descent allocates N GPU replicas to compute effectively identical updates -- a practice that leaves the rich s…
cs.LG2024
LiRank: Industrial Large Scale Ranking Models at LinkedIn
Fedor Borisyuk, Mingzhou Zhou, Qingquan Song +31
We present LiRank, a large-scale ranking framework at LinkedIn that brings to production state-of-the-art modeling architectures and optimization methods. We unveil several modelin…