1 paper
Jianbing Dong, Jianbin Chang
Training Large Language Models (LLMs) typically involves a two-stage pipeline at the output layer: hidden states are projected into vocabulary logits via a linear transformation (l…