papers

Publications (22)

cs.CL2026

Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context

Fitsum Reda, John Kamalu, Roger Waleffe +3

cs.LG2026

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

NVIDIA, :, Amala Sanjay Deshmukh +204

cs.LG2026

LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts

Venmugil Elango, Nidhi Bhatia, Roger Waleffe +13

cs.CL2025

Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

NVIDIA, :, Aaron Blakeman +198

cs.LG2022

MariusGNN: Resource-Efficient Out-of-Core Training of Graph Neural Networks

Roger Waleffe, Jason Mohoney, Theodoros Rekatsinas +1

cs.CL2025

NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

NVIDIA, :, Aarti Basant +214

cs.LG2023

Repeated Random Sampling for Minimizing the Time-to-Accuracy of Learning

Patrik Okanovic, Roger Waleffe, Vasilis Mageirakos +5

cs.CL2026

Pretraining Large Language Models with NVFP4

NVIDIA, Felix Abecassis, Anjulie Agrusa +87

cs.LG2026

Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aakshita Chandiramani +544

cs.CL2025

NVIDIA Nemotron 3: Efficient and Open Intelligence

NVIDIA, :, Aaron Blakeman +356

cs.LG2020

Principal Component Networks: Parameter Reduction Early in Training

Roger Waleffe, Theodoros Rekatsinas

cs.CL2026

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aaron Blakeman +571

cs.CL2025

Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aaron Blakeman +311

cs.LG2025

NVIDIA Nemotron Nano V2 VL

NVIDIA, :, Amala Sanjay Deshmukh +121

cs.CL2025

Llama-Nemotron: Efficient Reasoning Models

Akhiad Bercovich, Itay Levy, Izik Golan +132

cs.LG2025

Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models

Wenqi Jiang, Marco Zeller, Roger Waleffe +2

cs.CV2026

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi +293

cs.LG2025

Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks

Roger Waleffe, Devesh Sarda, Jason Mohoney +3

cs.LG2025

MLKV: Efficiently Scaling up Large Embedding Model Training with Disk-based Key-Value Storage

Yongjun He, Roger Waleffe, Zhichao Han +8

cs.LG2025

GraphSnapShot: Caching Local Structure for Fast Graph Learning

Dong Liu, Roger Waleffe, Meng Jiang +1

cs.LG2024

An Empirical Study of Mamba-based Language Models

Roger Waleffe, Wonmin Byeon, Duncan Riach +13

cs.LG2021

Marius: Learning Massive Graph Embeddings on a Single Machine

Jason Mohoney, Roger Waleffe, Yiheng Xu +2