178 citations · 270 across the 11 of their papers we have counts for
22 papers
Calibrating Factual Knowledge in Pretrained Language Models
Qingxiu Dong, Damai Dai, Yifan Song +3
Previous literature has proved that Pretrained Language Models (PLMs) can store factual knowledge. However, we find that facts stored in the PLMs are not always correct. It motivat…
Contextual Representation Learning beyond Masked Language Modeling
Zhiyi Fu, Wangchunshu Zhou, Jingjing Xu +2
How do masked language models (MLMs) such as BERT learn contextual representations? In this work, we analyze the learning dynamics of MLMs. We find that MLMs adopt sampled embeddin…
Graph-based Multi-hop Reasoning for Long Text Generation
Liang Zhao, Jingjing Xu, Junyang Lin +3
Long text generation is an important but challenging task.The main problem lies in learning sentence-level semantic dependencies which traditional generative models often suffer fr…
MUSE: Parallel Multi-Scale Attention for Sequence to Sequence Learning
Guangxiang Zhao, Xu Sun, Jingjing Xu +2
In sequence to sequence learning, the self-attention mechanism proves to be highly effective, and achieves significant improvements in many tasks. However, the self-attention mecha…
Understanding and Improving Layer Normalization
Jingjing Xu, Xu Sun, Zhiyuan Zhang +2
Layer normalization (LayerNorm) is a technique to normalize the distributions of intermediate layers. It enables smoother gradients, faster training, and better generalization accu…
Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering
Shangwen Lv, Daya Guo, Jingjing Xu +7
Commonsense question answering aims to answer questions which require background knowledge that is not explicitly expressed in the question. The key challenge is how to obtain evid…