1 paper · 1 filter
Cheng Luo, Zefan Cai, Junjie Hu
Transformers propagate information across depth through a single additive residual stream: every sublayer reads only the most recent state. Attention residuals relax this by lettin…