collaborators

5 papers

eess.AS2026

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation

Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo +3

Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored. Crucially, there is no consensus on whether audio-langu…

cs.RO2026

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation

Wei-Cheng Tseng, Gashon Hussein, Yuzhu Dong +9

Evaluating generalist robot manipulation policies in the real world is expensive, slow, and difficult to scale. Action-conditioned video world models offer a scalable alternative b…

cs.CV2026

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi +293

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…

cs.SD2026

Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods

Xuanru Zhou, Yiwen Shao, Wei-Cheng Tseng +1

Current audio pre-training seeks to learn unified representations for broad audio understanding tasks, but it remains fragmented and is fundamentally bottlenecked by its reliance o…

eess.AS2025

Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding

Mingyue Huo, Wei-Cheng Tseng, Yiwen Shao +2

Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward…