Masked Language Model Scoring
arXiv:1910.14659 · doi:10.18653/v1/2020.acl-main.240
Abstract
Pretrained masked language models (MLMs) require finetuning for most NLP tasks. Instead, we evaluate MLMs out of the box via their pseudo-log-likelihood scores (PLLs), which are computed by masking tokens one by one. We show that PLLs outperform scores from autoregressive language models like GPT-2 in a variety of tasks. By rescoring ASR and NMT hypotheses, RoBERTa reduces an end-to-end LibriSpeech model's WER by 30% relative and adds up to +1.7 BLEU on state-of-the-art baselines for low-resource translation pairs, with further gains from domain adaptation. We attribute this success to PLL's unsupervised expression of linguistic acceptability without a left-to-right bias, greatly improving on scores from GPT-2 (+10 points on island effects, NPI licensing in BLiMP). One can finetune MLMs to give scores without masking, enabling computation in a single inference pass. In all, PLLs and their associated pseudo-perplexities (PPPLs) enable plug-and-play use of the growing number of pretrained MLMs; e.g., we use a single cross-lingual model to rescore translations in multiple languages. We release our library for language model scoring at https://github.com/awslabs/mlm-scoring.
ACL 2020 camera-ready (presented July 2020)
Cited by in corpus (13)
- TransPolymer: a Transformer-based language model for polymer property predictions
- Innovative Bert-based Reranking Language Models for Speech Recognition
- RescoreBERT: Discriminative Speech Recognition Rescoring with BERT
- Non-autoregressive Transformer-based End-to-end ASR using BERT
- Testing the limits of natural language models for predicting human language judgments
- Are discrete units necessary for Spoken Language Modeling?
- Towards Effective Paraphrasing for Information Disguise
- SSAAM: Sentiment Signal-based Asset Allocation Method with Causality Information
- Fairness Definitions in Language Models Explained
- COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
- Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
- Comparative Analysis of Transformer Models in Disaster Tweet Classification for Public Safety
- Learning Project-wise Subsequent Code Edits via Interleaving Neural-based Induction and Tool-based Deduction