5 papers
Exploring Audio Hallucination in Egocentric Video Understanding
Ashish Seth, Xinhao Mei, Changsheng Zhao +9
Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly when visual information is unstab…
EgoAVU: Egocentric Audio-Visual Understanding
Ashish Seth, Xinhao Mei, Changsheng Zhao +9
Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due…
SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
Xinhao Mei, Gael Le Lan, Haohe Liu +5
Contrastive language-audio pretraining (CLAP) has achieved notable success in learning semantically rich audio representations and is widely adopted for various audio-related tasks…
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Haohe Liu, Gael Le Lan, Xinhao Mei +7
Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, produc…
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
Gael Le Lan, Bowen Shi, Zhaoheng Ni +9
We introduce MelodyFlow, an efficient text-controllable high-fidelity music generation and editing model. It operates on continuous latent representations from a low frame rate 48…