64 citations · 69 across the 4 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2022★ 64 cited
Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding Systems
Jack FitzGerald, Shankar Ananthakrishnan, Konstantine Arkoudas +38
We present results from a large-scale experiment on pretraining encoders with non-embedding parameter counts ranging from 700M to 9.3B, their subsequent distillation into smaller m…
cs.CL2021
Industry Scale Semi-Supervised Learning for Natural Language Understanding
Luoxin Chen, Francisco Garcia, Varun Kumar +2
This paper presents a production Semi-Supervised Learning (SSL) pipeline based on the student-teacher framework, which leverages millions of unlabeled examples to improve Natural L…