8 citations · 22 across the 9 of their papers we have counts for
11 papers
Building Synthetic Speaker Profiles in Text-to-Speech Systems
Jie Pu, Yixiong Meng, Oguz Elibol
The diversity of speaker profiles in multi-speaker TTS systems is a crucial aspect of its performance, as it measures how many different speaker profiles TTS systems could possibly…
Fusion of Embeddings Networks for Robust Combination of Text Dependent and Independent Speaker Recognition
Ruirui Li, Chelsea J. -T. Ju, Zeya Chen +3
By implicitly recognizing a user based on his/her speech input, speaker identification enables many downstream applications, such as personalized system behavior and expedited shop…
Scaling Laws for Acoustic Models
Jasha Droppo, Oguz Elibol
There is a recent trend in machine learning to increase model quality by growing models to sizes previously thought to be unreasonable. Recent work has shown that autoregressive ge…
Non-local convolutional neural networks (nlcnn) for speaker recognition
Haici Yang, Hongda Mao, Ruirui Li +2
Speaker recognition is the process of identifying a speaker based on the voice. The technology has attracted more attention with the recent increase in popularity of smart voice as…
Untangling in Invariant Speech Recognition
Cory Stephenson, Jenelle Feather, Suchismita Padhy +4
Encouraged by the success of deep neural networks on a variety of visual tasks, much theoretical and experimental work has been aimed at understanding and interpreting how vision n…
Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural Networks
Léopold Cambier, Anahita Bhiwandiwalla, Ting Gong +3
Training with larger number of parameters while keeping fast iterations is an increasingly adopted strategy and trend for developing better performing Deep Neural Network (DNN) mod…