collaborators

44 papers

cs.CV2026

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

Shufan Li, Greg Heinrich, Hanrong Ye +4

The paper introduces Nemotron-Labs-Diffusion-Image, a masked discrete diffusion model for high‑resolution text‑to‑image synthesis that adds a token‑editing mechanism and a grouped…

cs.CL2026

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

Yonggan Fu, Lexington Whalen, Abhinav Garg +23

We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architecture. Trained with a joint AR…

cs.CV2026

RADIO1D: Elastic Representations for Condensed Vision Modeling

Greg Heinrich, Mike Ranzinger, Collin McCarthy +6

This paper challenges the assumption that vision-language models (VLMs) require fixed patch-based 2D vision features. Analyzing fine-tuned vision encoders, we find that representat…

cs.AI2026

PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding

Zheng Wang, Zhifan Ye, Qi Cheng +8

Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, generating tokens in parallel. This makes them effective draft models f…

cs.CV2026

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

Yatai Ji, An-Chieh Cheng, Yang Fu +13

Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains…

cs.CV2026

Grounded 3D-Aware Spatial Vision-Language Modeling

An-Chieh Cheng, Yang Fu, Yatai Ji +12

We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding-…