1 citations · 2 across the 4 of their papers we have counts for
4 papers
On The Open Prompt Challenge In Conditional Audio Generation
Ernie Chang, Sidd Srinivasan, Mahi Luthra +8
Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is ch…
FoleyGen: Visually-Guided Audio Generation
Xinhao Mei, Varun Nagaraja, Gael Le Lan +4
Recent advancements in audio generation have been spurred by the evolution of large-scale deep learning models and expansive datasets. However, the task of video-to-audio (V2A) gen…
Stack-and-Delay: a new codebook pattern for music generation
Gael Le Lan, Varun Nagaraja, Ernie Chang +5
In language modeling based music generation, a generated waveform is represented by a sequence of hierarchical token stacks that can be decoded either in an auto-regressive manner…
Enhance audio generation controllability through representation similarity regularization
Yangyang Shi, Gael Le Lan, Varun Nagaraja +6
This paper presents an innovative approach to enhance control over audio generation by emphasizing the alignment between audio and text representations during model training. In th…