◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Mostofa Patwary

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL3

identity via Semantic Scholar / OpenAlex

most citedReuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models

5 citations · 7 across the 3 of their papers we have counts for

collaborators

3 papers

cs.CL2024★ 5 cited

Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models

Jupinder Parmar, Sanjev Satheesh, Mostofa Patwary +2

As language models have scaled both their number of parameters and pretraining dataset sizes, the computational cost for pretraining has become intractable except for the most well…

cs.CL2024

Nemotron-4 15B Technical Report

Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings +24

We introduce Nemotron-4 15B, a 15-billion-parameter large multilingual language model trained on 8 trillion text tokens. Nemotron-4 15B demonstrates strong performance when assesse…

cs.CL2023★ 2 cited

Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models

Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi +1

Pretrained large language models have become indispensable for solving various natural language processing (NLP) tasks. However, safely deploying them in real world applications is…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.