DiPair: Fast and Accurate Distillation for Trillion-Scale Text Matching and Pair Modeling
arXiv:2010.03099 · doi:10.18653/v1/2020.findings-emnlp.264
Abstract
Pre-trained models like BERT (Devlin et al., 2018) have dominated NLP / IR applications such as single sentence classification, text pair classification, and question answering. However, deploying these models in real systems is highly non-trivial due to their exorbitant computational costs. A common remedy to this is knowledge distillation (Hinton et al., 2015), leading to faster inference. However -- as we show here -- existing works are not optimized for dealing with pairs (or tuples) of texts. Consequently, they are either not scalable or demonstrate subpar performance. In this work, we propose DiPair -- a novel framework for distilling fast and accurate models on text pair tasks. Coupled with an end-to-end training strategy, DiPair is both highly scalable and offers improved quality-speed tradeoffs. Empirical studies conducted on both academic and real-world e-commerce benchmarks demonstrate the efficacy of the proposed approach with speedups of over 350x and minimal quality drop relative to the cross-attention teacher BERT model.
13 pages. Accepted to Findings of EMNLP 2020
References in corpus (7)
- Distilling the Knowledge in a Neural Network
- Convolutional Neural Network Architectures for Matching Natural Language Sentences
- A Deep Relevance Matching Model for Ad-hoc Retrieval
- End-to-End Neural Ad-hoc Ranking with Kernel Pooling
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Comparing Rewinding and Fine-tuning in Neural Network Pruning
- To Tune or Not To Tune? How About the Best of Both Worlds?
Cited by in corpus (5)
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Composite Re-Ranking for Efficient Document Search with BERT
- Condenser: a Pre-training Architecture for Dense Retrieval
- Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations
- Evaluation of Semantic Answer Similarity Metrics