1 paper
Peng Xu, Xinchi Chen, Xiaofei Ma +2
Recent progress in pretrained Transformer-based language models has shown great success in learning contextual representation of text. However, due to the quadratic self-attention…