1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Shyam Venkatasubramanian, Sean Moushegian, Michael Lin +3
Standard attention-based transformers are known to exhibit instability under learning rate overspecification during training, particularly at high learning rates. While various met…