8 citations · 19 across the 22 of their papers we have counts for
4 papers · 1 filter
Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
Susan Liang, Chao Huang, Yapeng Tian +2
In this paper, we introduce a novel task called language-guided joint audio-visual editing. Given an audio and image pair of a sounding event, this task aims at generating new audi…
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
Yolo Y. Tang, Junjia Guo, Hang Hua +9
The advancement of Multimodal Large Language Models (MLLMs) has enabled significant progress in multimodal understanding, expanding their capacity to analyze video content. However…
Scaling Concept With Text-Guided Diffusion Models
Chao Huang, Susan Liang, Yunlong Tang +3
Text-guided diffusion models have revolutionized generative tasks by producing high-fidelity content from text descriptions. They have also enabled an editing paradigm where concep…
Modeling and Driving Human Body Soundfields through Acoustic Primitives
Chao Huang, Dejan Markovic, Chenliang Xu +1
While rendering and animation of photorealistic 3D human body models have matured and reached an impressive quality over the past years, modeling the spatial audio associated with…