collaborators

51 papers

eess.AS2026

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar +19

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, a…

cs.LG2026

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Mikail Khona, Aditya Vavre, Boxiang Wang +11

Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale. I…

cs.CL2026

Unified Audio Intelligence Without Regressing on Text Intelligence

Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim +17

Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-te…

cs.LG2026

FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data

Tao Feng, Haozhen Zhang, Zijie Lei +5

The rapid advancement of large language models (LLMs) has created a diverse landscape of models, each excelling at different tasks. This diversity drives researchers to employ mult…

cs.CL2026

Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context

Fitsum Reda, John Kamalu, Roger Waleffe +3

Diffusion language models offer a promising alternative to autoregressive models due to their potential for parallel and iterative generation. However, existing approaches use a si…

cs.CL2026

MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos

Arushi Goel, Sreyan Ghosh, Vatsal Agarwal +16

Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over…