activity
20182021
most citedQuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions

31 citations · 70 across the 3 of their papers we have counts for

collaborators

11 papers

cs.CL202119 cited

Offensive Language and Hate Speech Detection with Deep Learning and Transfer Learning

Bencheng Wei, Jason Li, Ajay Gupta +3

Toxic online speech has become a crucial problem nowadays due to an exponential increase in the use of internet by people from different cultures and educational backgrounds. Diffe…

eess.AS2020

SpeakerNet: 1D Depth-wise Separable Convolutional Network for Text-Independent Speaker Recognition and Verification

Nithin Rao Koluguri, Jason Li, Vitaly Lavrukhin +1

We propose SpeakerNet - a new neural architecture for speaker recognition and speaker verification tasks. It is composed of residual blocks with 1D depth-wise separable convolution…

eess.AS202020 cited

Cross-Language Transfer Learning, Continuous Learning, and Domain Adaptation for End-to-End Automatic Speech Recognition

Jocelyn Huang, Oleksii Kuchaiev, Patrick O'Neill +5

In this paper, we demonstrate the efficacy of transfer learning and continuous learning for various automatic speech recognition (ASR) tasks. We start with a pre-trained English AS…

cs.CV2020

Cycle Text-To-Image GAN with BERT

Trevor Tsue, Samir Sen, Jason Li

We explore novel approaches to the task of image generation from their respective captions, building on state-of-the-art GAN architectures. Particularly, we baseline our models wit…

cs.SD2019

Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens

Rafael Valle, Jason Li, Ryan Prenger +1

Mellotron is a multispeaker voice synthesis model based on Tacotron 2 GST that can make a voice emote and sing without emotive or singing training data. By explicitly conditioning…

eess.AS201931 cited

QuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions

Samuel Kriman, Stanislav Beliaev, Boris Ginsburg +6

We propose a new end-to-end neural acoustic model for automatic speech recognition. The model is composed of multiple blocks with residual connections between them. Each block cons…