1 paper
Ian Porada, Alessandro Sordoni, Jackie Chi Kit Cheung
Transformer models pre-trained with a masked-language-modeling objective (e.g., BERT) encode commonsense knowledge as evidenced by behavioral probes; however, the extent to which t…