1 paper
Xingjian Wang, Qingyu Han, Xiaodong Luo +1
Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate u…