17 citations · 17 across the 7 of their papers we have counts for
1 paper · 1 filter
Weijie Zhao, Mingquan Liu, Bolun Wang +4
Scaling Transformers typically necessitates training larger models from scratch, as standard architectures struggle to expand without discarding learned representations. We identif…