24 citations · 30 across the 2 of their papers we have counts for
3 papers
High-Fidelity Audio Compression with Improved RVQGAN
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs +2
Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model…
Wav2CLIP: Learning Robust Audio Representations From CLIP
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar +1
We propose Wav2CLIP, a robust audio representation learning method by distilling from Contrastive Language-Image Pre-training (CLIP). We systematically evaluate Wav2CLIP on a varie…
Chunked Autoregressive GAN for Conditional Waveform Synthesis
Max Morrison, Rithesh Kumar, Kundan Kumar +3
Conditional waveform synthesis models learn a distribution of audio waveforms given conditioning such as text, mel-spectrograms, or MIDI. These systems employ deep generative model…