19 papers
Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion
Dongseok Shim, Julian Tanke, Kengo Uchida +5
Human motion generation has been widely studied across diverse input modalities, text, music, and video, and recent efforts have unified these into single multimodal frameworks. Ho…
Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis
Shuyang Cui, Zhi Zhong, Qiyu Wu +9
Current methods for creating drum loop audio in digital music production, such as using one-shot samples or resampling, often demand non-trivial efforts of creators. While recent g…
MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video
Kazuya Tateishi, Akira Takahashi, Atsuo Hiroe +3
Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound production, demand not only the genera…
MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation
Akira Takahashi, Ryosuke Sawata, Shusuke Takahashi +1
Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model room-acoustic effects such as re…
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
Akira Takahashi, Shusuke Takahashi, Yuki Mitsufuji
We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By leveraging knowledge about the relation…
Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
Christian Simon, Masato Ishii, Wei-Yao Wang +8
Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information.…