1 citations · 5 across the 9 of their papers we have counts for
Showing 2024Show all
3 papers · 1 filter
cs.MM2024
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Haohe Liu, Gael Le Lan, Xinhao Mei +7
Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, produc…
eess.AS2024
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
Ruizhe Huang, Xiaohui Zhang, Zhaoheng Ni +9
Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can c…
eess.AS2024★ 1 cited
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
Gael Le Lan, Bowen Shi, Zhaoheng Ni +9
We introduce MelodyFlow, an efficient text-controllable high-fidelity music generation and editing model. It operates on continuous latent representations from a low frame rate 48…