5 papers
MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs
Daeyong Kwon, Qiyu Wu, Shinobu Kuriya +6
Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses are grounded in the correct temp…
Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis
Shuyang Cui, Zhi Zhong, Qiyu Wu +9
Current methods for creating drum loop audio in digital music production, such as using one-shot samples or resampling, often demand non-trivial efforts of creators. While recent g…
Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
Christian Simon, Masato Ishii, Wei-Yao Wang +8
Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information.…
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
Qiyu Wu, Shuyang Cui, Satoshi Hayakawa +3
Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the suc…
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
Zhi Zhong, Akira Takahashi, Shuyang Cui +3
Foley synthesis aims to synthesize high-quality audio that is both semantically and temporally aligned with video frames. Given its broad application in creative industries, the ta…