collaborators

5 papers

cs.CV2026

Exploring Audio Hallucination in Egocentric Video Understanding

Ashish Seth, Xinhao Mei, Changsheng Zhao +9

Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly when visual information is unstab…

cs.CV2026

EgoAVU: Egocentric Audio-Visual Understanding

Ashish Seth, Xinhao Mei, Changsheng Zhao +9

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due…

eess.AS2026

SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training

Xinhao Mei, Gael Le Lan, Haohe Liu +5

Contrastive language-audio pretraining (CLAP) has achieved notable success in learning semantically rich audio representations and is widely adopted for various audio-related tasks…

cs.MM2024

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Haohe Liu, Gael Le Lan, Xinhao Mei +7

Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, produc…

eess.AS2024

High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

Gael Le Lan, Bowen Shi, Zhaoheng Ni +9

We introduce MelodyFlow, an efficient text-controllable high-fidelity music generation and editing model. It operates on continuous latent representations from a low frame rate 48…