papers

Publications (10)

cs.CL2021

Text Mining Drug/Chemical-Protein Interactions using an Ensemble of BERT and T5 Based Models

Virginia Adams, Hoo-Chang Shin, Carol Anderson +2

In Track-1 of the BioCreative VII Challenge participants are asked to identify interactions between drugs/chemicals and proteins. In-context named entity annotations for each drug/…

cs.LG2026

Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aakshita Chandiramani +544

We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention Mixture-of-Experts model. Nemo…

cs.CL2021

NVIDIA NeMo Neural Machine Translation Systems for English-German and English-Russian News and Biomedical Tasks at WMT21

Sandeep Subramanian, Oleksii Hrinchuk, Virginia Adams +1

This paper provides an overview of NVIDIA NeMo's neural machine translation systems for the constrained data track of the WMT21 News and Biomedical Shared Translation Tasks. Our ne…

cs.CL2022

Finding the Right Recipe for Low Resource Domain Adaptation in Neural Machine Translation

Virginia Adams, Sandeep Subramanian, Mike Chrzanowski +2

General translation models often still struggle to generate accurate translations in specialized domains. To guide machine translation practitioners and characterize the effectiven…

cs.CL2024

RedPajama: an Open Dataset for Training Large Language Models

Maurice Weber, Daniel Fu, Quentin Anthony +16

Large language models are increasingly becoming a cornerstone technology in artificial intelligence, the sciences, and society as a whole, yet the optimal strategies for dataset co…

cs.CL2023

HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM

Zhilin Wang, Yi Dong, Jiaqi Zeng +8

Existing open-source helpfulness preference datasets do not specify what makes some responses more helpful and others less so. Models trained on these datasets can incidentally lea…