1 citations · 1 across the 2 of their papers we have counts for
7 papers
Do Joint Audio-Video Generation Models Understand Physics?
Zijun Cui, Xiulong Liu, Hao Fang +8
Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visual physics, or merely generate…
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
Weiguo Pian, Shijian Deng, Shentong Mo +3
In this paper, we introduce Modality-Inconsistent Continual Learning (MICL), a new continual learning scenario for Multimodal Large Language Models (MLLMs) that involves tasks with…
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
Weiguo Pian, Saksham Singh Kushwaha, Zhimin Chen +4
In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds ac…
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
Yue Zhang, Liqiang Jing, Jia Li +4
Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approa…
From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
Jia Li, Yapeng Tian
Audio-Visual Segmentation (AVS) aims to identify and segment sound-producing objects in videos by leveraging both visual and audio modalities. It has emerged as a significant resea…
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
Jia Li, Wenjie Zhao, Ziru Huang +2
Unlike traditional visual segmentation, audio-visual segmentation (AVS) requires the model not only to identify and segment objects but also to determine whether they are sound sou…