2 citations · 5 across the 6 of their papers we have counts for
6 papers
Masked Audio Generation using a Single Non-Autoregressive Transformer
Alon Ziv, Itai Gat, Gael Le Lan +6
We introduce MAGNeT, a masked generative sequence modeling method that operates directly over several streams of audio tokens. Unlike prior work, MAGNeT is comprised of a single-st…
In-Context Prompt Editing For Conditional Audio Generation
Ernie Chang, Pin-Jie Lin, Yang Li +6
Distributional shift is a central challenge in the deployment of machine learning models as they can be ill-equipped for real-world data. This is particularly evident in text-to-au…
Exploring Speech Enhancement for Low-resource Speech Synthesis
Zhaoheng Ni, Sravya Popuri, Ning Dong +6
High-quality and intelligible speech is essential to text-to-speech (TTS) model training, however, obtaining high-quality data for low-resource languages is challenging and expensi…
FoleyGen: Visually-Guided Audio Generation
Xinhao Mei, Varun Nagaraja, Gael Le Lan +4
Recent advancements in audio generation have been spurred by the evolution of large-scale deep learning models and expansive datasets. However, the task of video-to-audio (V2A) gen…
Stack-and-Delay: a new codebook pattern for music generation
Gael Le Lan, Varun Nagaraja, Ernie Chang +5
In language modeling based music generation, a generated waveform is represented by a sequence of hierarchical token stacks that can be decoded either in an auto-regressive manner…
Enhance audio generation controllability through representation similarity regularization
Yangyang Shi, Gael Le Lan, Varun Nagaraja +6
This paper presents an innovative approach to enhance control over audio generation by emphasizing the alignment between audio and text representations during model training. In th…