10 citations · 10 across the 3 of their papers we have counts for
3 papers
cs.CL2023
Understanding Shared Speech-Text Representations
Gary Wang, Kyle Kastner, Ankur Bapna +4
Recently, a number of approaches to train speech models by incorpo-rating text into end-to-end models have been developed, with Mae-stro advancing state-of-the-art automatic speech…
cs.SD2022
R-MelNet: Reduced Mel-Spectral Modeling for Neural TTS
Kyle Kastner, Aaron Courville
This paper introduces R-MelNet, a two-part autoregressive architecture with a frontend based on the first tier of MelNet and a backend WaveRNN-style audio decoder for neural text-t…
cs.SD2021★ 10 cited
MIDI-DDSP: Detailed Control of Musical Performance via Hierarchical Modeling
Yusong Wu, Ethan Manilow, Yi Deng +6
Musical expression requires control of both what notes are played, and how they are performed. Conventional audio synthesizers provide detailed expressive controls, but at the cost…