5 papers
RADIO1D: Elastic Representations for Condensed Vision Modeling
Greg Heinrich, Mike Ranzinger, Collin McCarthy +6
This paper challenges the assumption that vision-language models (VLMs) require fixed patch-based 2D vision features. Analyzing fine-tuned vision encoders, we find that representat…
Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders
Yitong Jiang, Hongjun Wang, Collin McCarthy +15
Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale pretraining. Subquadratic al…
Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
NVIDIA, :, Amala Sanjay Deshmukh +204
We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 N…
C-RADIOv4 (Tech Report)
Mike Ranzinger, Greg Heinrich, Collin McCarthy +4
By leveraging multi-teacher distillation, agglomerative vision backbones provide a unified student model that retains and improves the distinct capabilities of multiple teachers. I…
GSPN-2: Efficient Parallel Sequence Modeling
Hongjun Wang, Yitong Jiang, Collin McCarthy +12
Efficient vision transformer remains a bottleneck for high-resolution images and long-video related real-world applications. Generalized Spatial Propagation Network (GSPN) addresse…