3 papers
cs.CV2026
VLM3: Vision Language Models Are Native 3D Learners
Zhipeng Cai, Zhuang Liu, Yunyang Xiong +3
Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D u…
eess.AS2026
SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
Xinhao Mei, Gael Le Lan, Haohe Liu +5
Contrastive language-audio pretraining (CLAP) has achieved notable success in learning semantically rich audio representations and is widely adopted for various audio-related tasks…
cs.MM2024
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Haohe Liu, Gael Le Lan, Xinhao Mei +7
Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, produc…