1 paper
Qingyu Yin, Xuzheng He, Xiang Zhuang +4
The decoder-only Transformer architecture with causal masking and relative position encoding (RPE) has become the de facto choice in language modeling. Despite its exceptional perf…