4 papers
Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
Christian Simon, Masato Ishii, Wei-Yao Wang +8
Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information.…
SF-Mamba: Rethinking State Space Model for Vision
Masakazu Yoshimura, Teruaki Hayashi, Yuki Hoshino +2
The realm of Mamba for vision has been advanced in recent years to strike for the alternatives of Vision Transformers (ViTs) that suffer from the quadratic complexity. While the re…
VIRTUE: Visual-Interactive Text-Image Universal Embedder
Wei-Yao Wang, Kazuya Tateishi, Qiyu Wu +2
Multimodal representation learning models have demonstrated successful operation across complex tasks, and the integration of vision-language models (VLMs) has further enabled embe…
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
Qiyu Wu, Shuyang Cui, Satoshi Hayakawa +3
Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the suc…