14 citations · 40 across the 12 of their papers we have counts for
5 papers · 1 filter
Neural Pitch-Shifting and Time-Stretching with Controllable LPCNet
Max Morrison, Zeyu Jin, Nicholas J. Bryan +2
Modifying the pitch and timing of an audio signal are fundamental audio editing operations with applications in speech manipulation, audio-visual synchronization, and singing voice…
Context-Aware Prosody Correction for Text-Based Speech Editing
Max Morrison, Lucas Rencker, Zeyu Jin +3
Text-based speech editors expedite the process of editing speech recordings by permitting editing via intuitive cut, copy, and paste operations on a speech transcript. A major draw…
AutoClip: Adaptive Gradient Clipping for Source Separation Networks
Prem Seetharaman, Gordon Wichern, Bryan Pardo +1
Clipping the gradient is a known approach to improving gradient descent, but requires hand selection of a clipping threshold hyperparameter. We present AutoClip, a simple method fo…
Model selection for deep audio source separation via clustering analysis
Alisa Liu, Prem Seetharaman, Bryan Pardo
Audio source separation is the process of separating a mixture (e.g. a pop band recording) into isolated sounds from individual sources (e.g. just the lead vocals). Deep learning m…
Simultaneous Separation and Transcription of Mixtures with Multiple Polyphonic and Percussive Instruments
Ethan Manilow, Prem Seetharaman, Bryan Pardo
We present a single deep learning architecture that can both separate an audio recording of a musical mixture into constituent single-instrument recordings and transcribe these ins…