1 paper
Jie Zhang, Mao-Hsuan Mao, Bo-Wei Chiu +1
Recent advances in deep learning have established Transformer architectures as the predominant modeling paradigm. Central to the success of Transformers is the self-attention mecha…