1 paper
Onkar Pandit, Yufang Hou
We probe pre-trained transformer language models for bridging inference. We first investigate individual attention heads in BERT and observe that attention heads at higher layers p…