activity
20172021
most citedFlowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis

81 citations · 89 across the 4 of their papers we have counts for

collaborators

15 papers

cs.SD2021

One TTS Alignment To Rule Them All

Rohan Badlani, Adrian Łancucki, Kevin J. Shih +3

Speech-to-text alignment is a critical component of neural textto-speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-l…

cs.CL20218 cited

Clinical Named Entity Recognition using Contextualized Token Representations

Yichao Zhou, Chelsea Ju, J. Harry Caufield +6

The clinical named entity recognition (CNER) task seeks to locate and classify clinical terminologies into predefined categories, such as diagnostic procedure, disease disorder, se…

cs.SD202081 cited

Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis

Rafael Valle, Kevin Shih, Ryan Prenger +1

In this paper we propose Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis with control over speech variation and style transfer. Flowtron borr…

cs.CV2020

Unsupervised Disentanglement of Pose, Appearance and Background from Images and Videos

Aysegul Dundar, Kevin J. Shih, Animesh Garg +3

Unsupervised landmark learning is the task of learning semantic keypoint-like representations without the use of expensive input keypoint-level annotations. A popular approach is t…

cs.CV2019

Video Interpolation and Prediction with Unsupervised Landmarks

Kevin J. Shih, Aysegul Dundar, Animesh Garg +3

Prediction and interpolation for long-range video data involves the complex task of modeling motion trajectories for each visible object, occlusions and dis-occlusions, as well as…

cs.CV2019

Unsupervised Video Interpolation Using Cycle Consistency

Fitsum A. Reda, Deqing Sun, Aysegul Dundar +6

Learning to synthesize high frame rate videos via interpolation requires large quantities of high frame rate training videos, which, however, are scarce, especially at high resolut…