8 citations · 13 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024★ 1 cited
Scaling Exponents Across Parameterizations and Optimizers
Katie Everett, Lechao Xiao, Mitchell Wortsman +8
Robust and effective scaling of models from small to large width typically requires the precise adjustment of many algorithmic and architectural details, such as parameterization a…
cs.LG2023★ 4 cited
Small-scale proxies for large-scale Transformer training instabilities
Mitchell Wortsman, Peter J. Liu, Lechao Xiao +13
Teams that have trained large Transformer-based models have reported training instabilities at large scale that did not appear when training with the same hyperparameters at smalle…