The MultiBERTs: BERT Reproductions for Robustness Analysis
arXiv:2106.16163
Abstract
Experiments with pre-trained models such as BERT are often based on a single checkpoint. While the conclusions drawn apply to the artifact tested in the experiment (i.e., the particular instance of the model), it is not always clear whether they hold for the more general procedure which includes the architecture, training data, initialization scheme, and loss function. Recent work has shown that repeating the pre-training process can lead to substantially different performance, suggesting that an alternate strategy is needed to make principled statements about procedures. To enable researchers to draw more robust conclusions, we introduce the MultiBERTs, a set of 25 BERT-Base checkpoints, trained with similar hyper-parameters as the original BERT model but differing in random weight initialization and shuffling of training data. We also define the Multi-Bootstrap, a non-parametric bootstrap method for statistical inference designed for settings where there are multiple pre-trained models and limited test data. To illustrate our approach, we present a case study of gender bias in coreference resolution, in which the Multi-Bootstrap lets us measure effects that may not be detected with a single checkpoint. We release our models and statistical library along with an additional set of 140 intermediate checkpoints captured during pre-training to facilitate research on learning dynamics.
Accepted at ICLR'22. Checkpoints and example analyses: http://goo.gle/multiberts
References in corpus (6)
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
- Quantifying the Carbon Emissions of Machine Learning
- Carbon Emissions and Large Neural Network Training
- Are Larger Pretrained Language Models Uniformly Better? Comparing Performance at the Instance Level
Cited by in corpus (7)
- A Multi-Center Study on the Adaptability of a Shared Foundation Model for Electronic Health Records
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- On The Impact of Machine Learning Randomness on Group Fairness
- Shatter: An Efficient Transformer Encoder with Single-Headed Self-Attention and Relative Sequence Partitioning
- All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality
- Incorporating Residual and Normalization Layers into Analysis of Masked Language Models
- How Emotionally Stable is ALBERT? Testing Robustness with Stochastic Weight Averaging on a Sentiment Analysis Task