167 citations · 216 across the 15 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2021★ 1 cited
Convolutions and Self-Attention: Re-interpreting Relative Positions in Pre-trained Language Models
Tyler A. Chang, Yifan Xu, Weijian Xu +1
In this paper, we detail the relationship between convolutions and self-attention in natural language tasks. We show that relative position embeddings in self-attention layers are…
cs.CL2019
Rethinking Exposure Bias In Language Modeling
Yifan Xu, Kening Zhang, Haoyu Dong +3
Exposure bias describes the phenomenon that a language model trained under the teacher forcing schema may perform poorly at the inference stage when its predictions are conditioned…