most citedMasked Audio Generation using a Single Non-Autoregressive Transformer

2 citations · 5 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD20242 cited

Masked Audio Generation using a Single Non-Autoregressive Transformer

Alon Ziv, Itai Gat, Gael Le Lan +6

We introduce MAGNeT, a masked generative sequence modeling method that operates directly over several streams of audio tokens. Unlike prior work, MAGNeT is comprised of a single-st…

cs.SD20231 cited

In-Context Prompt Editing For Conditional Audio Generation

Ernie Chang, Pin-Jie Lin, Yang Li +6

Distributional shift is a central challenge in the deployment of machine learning models as they can be ill-equipped for real-world data. This is particularly evident in text-to-au…

eess.AS20231 cited

Exploring Speech Enhancement for Low-resource Speech Synthesis

Zhaoheng Ni, Sravya Popuri, Ning Dong +6

High-quality and intelligible speech is essential to text-to-speech (TTS) model training, however, obtaining high-quality data for low-resource languages is challenging and expensi…

eess.AS20231 cited

FoleyGen: Visually-Guided Audio Generation

Xinhao Mei, Varun Nagaraja, Gael Le Lan +4

Recent advancements in audio generation have been spurred by the evolution of large-scale deep learning models and expansive datasets. However, the task of video-to-audio (V2A) gen…

eess.AS2023

Stack-and-Delay: a new codebook pattern for music generation

Gael Le Lan, Varun Nagaraja, Ernie Chang +5

In language modeling based music generation, a generated waveform is represented by a sequence of hierarchical token stacks that can be decoded either in an auto-regressive manner…

cs.SD2023

Enhance audio generation controllability through representation similarity regularization

Yangyang Shi, Gael Le Lan, Varun Nagaraja +6

This paper presents an innovative approach to enhance control over audio generation by emphasizing the alignment between audio and text representations during model training. In th…