6 citations · 10 across the 3 of their papers we have counts for
1 paper · 1 filter
Wuyang Chen, Yanqi Zhou, Nan Du +4
Pretraining on a large-scale corpus has become a standard method to build general language models (LMs). Adapting a model to new data distributions targeting different downstream t…