81 citations · 89 across the 4 of their papers we have counts for
15 papers
One TTS Alignment To Rule Them All
Rohan Badlani, Adrian Łancucki, Kevin J. Shih +3
Speech-to-text alignment is a critical component of neural textto-speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-l…
Clinical Named Entity Recognition using Contextualized Token Representations
Yichao Zhou, Chelsea Ju, J. Harry Caufield +6
The clinical named entity recognition (CNER) task seeks to locate and classify clinical terminologies into predefined categories, such as diagnostic procedure, disease disorder, se…
Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
Rafael Valle, Kevin Shih, Ryan Prenger +1
In this paper we propose Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis with control over speech variation and style transfer. Flowtron borr…
Unsupervised Disentanglement of Pose, Appearance and Background from Images and Videos
Aysegul Dundar, Kevin J. Shih, Animesh Garg +3
Unsupervised landmark learning is the task of learning semantic keypoint-like representations without the use of expensive input keypoint-level annotations. A popular approach is t…
Video Interpolation and Prediction with Unsupervised Landmarks
Kevin J. Shih, Aysegul Dundar, Animesh Garg +3
Prediction and interpolation for long-range video data involves the complex task of modeling motion trajectories for each visible object, occlusions and dis-occlusions, as well as…
Unsupervised Video Interpolation Using Cycle Consistency
Fitsum A. Reda, Deqing Sun, Aysegul Dundar +6
Learning to synthesize high frame rate videos via interpolation requires large quantities of high frame rate training videos, which, however, are scarce, especially at high resolut…