5 citations · 7 across the 3 of their papers we have counts for
3 papers
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
Jupinder Parmar, Sanjev Satheesh, Mostofa Patwary +2
As language models have scaled both their number of parameters and pretraining dataset sizes, the computational cost for pretraining has become intractable except for the most well…
Nemotron-4 15B Technical Report
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings +24
We introduce Nemotron-4 15B, a 15-billion-parameter large multilingual language model trained on 8 trillion text tokens. Nemotron-4 15B demonstrates strong performance when assesse…
Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models
Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi +1
Pretrained large language models have become indispensable for solving various natural language processing (NLP) tasks. However, safely deploying them in real world applications is…