3 papers
cs.CV2026
Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
Christian Simon, Masato Ishii, Wei-Yao Wang +8
Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information.…
cs.CL2025
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
Qiyu Wu, Shuyang Cui, Satoshi Hayakawa +3
Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the suc…
cs.SD2025
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
Zhi Zhong, Akira Takahashi, Shuyang Cui +3
Foley synthesis aims to synthesize high-quality audio that is both semantically and temporally aligned with video frames. Given its broad application in creative industries, the ta…