72 citations · 83 across the 2 of their papers we have counts for
5 papers
Fully Supervised Speaker Diarization
Aonan Zhang, Quan Wang, Zhenyao Zhu +2
In this paper, we propose a fully supervised speaker diarization approach, named unbounded interleaved-state recurrent neural networks (UIS-RNN). Given extracted speaker-discrimina…
Exploring Neural Transducers for End-to-End Speech Recognition
Eric Battenberg, Jitong Chen, Rewon Child +8
In this work, we perform an empirical comparison among the CTC, RNN-Transducer, and attention-based Seq2Seq models for end-to-end speech recognition. We show that, without any lang…
Reducing Bias in Production Speech Models
Eric Battenberg, Rewon Child, Adam Coates +13
Replacing hand-engineered pipelines with end-to-end deep learning systems has enabled strong results in applications like speech and object recognition. However, the causality and…
Deep Speaker: an End-to-End Neural Speaker Embedding System
Chao Li, Xiaokong Ma, Bing Jiang +6
We present Deep Speaker, a neural speaker embedding system that maps utterances to a hypersphere where speaker similarity is measured by cosine similarity. The embeddings generated…
Learning Multiscale Features Directly From Waveforms
Zhenyao Zhu, Jesse H. Engel, Awni Hannun
Deep learning has dramatically improved the performance of speech recognition systems through learning hierarchies of features optimized for the task at hand. However, true end-to-…