67 citations · 90 across the 16 of their papers we have counts for
3 papers · 1 filter
Exploring Speech Enhancement for Low-resource Speech Synthesis
Zhaoheng Ni, Sravya Popuri, Ning Dong +6
High-quality and intelligible speech is essential to text-to-speech (TTS) model training, however, obtaining high-quality data for low-resource languages is challenging and expensi…
FoleyGen: Visually-Guided Audio Generation
Xinhao Mei, Varun Nagaraja, Gael Le Lan +4
Recent advancements in audio generation have been spurred by the evolution of large-scale deep learning models and expansive datasets. However, the task of video-to-audio (V2A) gen…
Stack-and-Delay: a new codebook pattern for music generation
Gael Le Lan, Varun Nagaraja, Ernie Chang +5
In language modeling based music generation, a generated waveform is represented by a sequence of hierarchical token stacks that can be decoded either in an auto-regressive manner…