1 citations · 4 across the 8 of their papers we have counts for
10 papers
SongCraft: Unified Song Generation and Editing with Reconstructive Learning
Haohe Liu, Varun Nagaraja, Gael Le Lan +5
Song generation and editing have mostly been treated as separate tasks. Existing editing methods often require noise injection and regeneration or curated paired training data. We…
SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
Xinhao Mei, Gael Le Lan, Haohe Liu +5
Contrastive language-audio pretraining (CLAP) has achieved notable success in learning semantically rich audio representations and is widely adopted for various audio-related tasks…
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Haohe Liu, Gael Le Lan, Xinhao Mei +7
Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, produc…
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
Ruizhe Huang, Xiaohui Zhang, Zhaoheng Ni +9
Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can c…
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
Gael Le Lan, Bowen Shi, Zhaoheng Ni +9
We introduce MelodyFlow, an efficient text-controllable high-fidelity music generation and editing model. It operates on continuous latent representations from a low frame rate 48…
On The Open Prompt Challenge In Conditional Audio Generation
Ernie Chang, Sidd Srinivasan, Mahi Luthra +8
Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is ch…