4 citations · 5 across the 3 of their papers we have counts for
3 papers
cs.LG2026
Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth
Katie Everett
We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip con…
cs.LG2024★ 1 cited
Scaling Exponents Across Parameterizations and Optimizers
Katie Everett, Lechao Xiao, Mitchell Wortsman +8
Robust and effective scaling of models from small to large width typically requires the precise adjustment of many algorithmic and architectural details, such as parameterization a…
cs.LG2023★ 4 cited
Small-scale proxies for large-scale Transformer training instabilities
Mitchell Wortsman, Peter J. Liu, Lechao Xiao +13
Teams that have trained large Transformer-based models have reported training instabilities at large scale that did not appear when training with the same hyperparameters at smalle…