collaborators

43 papers

eess.AS2026

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar +19

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, a…

cs.LG2026

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Mikail Khona, Aditya Vavre, Boxiang Wang +11

Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale. I…

cs.CL2026

Unified Audio Intelligence Without Regressing on Text Intelligence

Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim +17

Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-te…

cs.LG2026

FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data

Tao Feng, Haozhen Zhang, Zijie Lei +5

The rapid advancement of large language models (LLMs) has created a diverse landscape of models, each excelling at different tasks. This diversity drives researchers to employ mult…

cs.CL2026

Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context

Fitsum Reda, John Kamalu, Roger Waleffe +3

Diffusion language models offer a promising alternative to autoregressive models due to their potential for parallel and iterative generation. However, existing approaches use a si…

cs.CV2026

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi +293

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…