6.7k citations · 19.4k across the 18 of their papers we have counts for
Showing eess.ASShow all
2 papers · 1 filter
eess.AS2022★ 1.2k cited
Robust Speech Recognition via Large-Scale Weak Supervision
Alec Radford, Jong Wook Kim, Tao Xu +3
We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual…
eess.AS2020★ 107 cited
Jukebox: A Generative Model for Music
Prafulla Dhariwal, Heewoo Jun, Christine Payne +3
We introduce Jukebox, a model that generates music with singing in the raw audio domain. We tackle the long context of raw audio using a multi-scale VQ-VAE to compress it to discre…