activity
20242026
most citedAudio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing

1 citations · 2 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD2026★ 1 cited

Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing

Zeyue Tian, Binxin Yang, Zhaoyang Liu +8

Recent progress in multimodal models has spurred rapid advances in audio understanding, generation, and editing. However, these capabilities are typically addressed by specialized…

cs.MA2026

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

Bowen Yang, Kaiming Jin, Zhenyu Wu +12

While Vision-Language Models (VLMs) have significantly advanced Computer-Using Agents (CUAs), current frameworks struggle with robustness in long-horizon workflows and generalizati…

cs.LG2025

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems

Brendan Leigh Ross, Noël Vouitsis, Atiyeh Ashari Ghomi +8

Although large language models (LLMs) are becoming increasingly capable of solving challenging real-world tasks, accurately quantifying their uncertainty remains a critical open pr…

cs.MM2025

AudioX: A Unified Framework for Anything-to-Audio Generation

Zeyue Tian, Zhaoyang Liu, Yizhu Jin +6

Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a unified multimodal modeling framework,…

cs.CV2024

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement

Zhefan Rao, Liya Ji, Yazhou Xing +6

Text-to-video (T2V) generation has gained significant attention recently. However, the costs of training a T2V model from scratch remain persistently high, and there is considerabl…

cs.CV2024★ 1 cited

MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions

Xiaowei Chi, Yatian Wang, Aosong Cheng +16

Massive multi-modality datasets play a significant role in facilitating the success of large video-language models. However, current video-language datasets primarily provide text…