activity
20172024
most citedGigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

216 citations · 259 across the 27 of their papers we have counts for

collaborators
Showing 2020 · eess.ASShow all

8 papers · 2 filters

eess.AS2020

DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs

Desh Raj, Leibny Paola Garcia-Perera, Zili Huang +4

Several advances have been made recently towards handling overlapping speech for speaker diarization. Since speech and natural language tasks often benefit from ensemble techniques…

eess.AS2020★ 4 cited

Frustratingly Easy Noise-aware Training of Acoustic Models

Desh Raj, Jesus Villalba, Daniel Povey +1

Environmental noises and reverberation have a detrimental effect on the performance of automatic speech recognition (ASR) systems. Multi-condition training of neural network-based…

eess.AS2020

Neural Language Modeling With Implicit Cache Pointers

Ke Li, Daniel Povey, Sanjeev Khudanpur

A cache-inspired approach is proposed for neural language models (LMs) to improve long-range dependency and better predict rare words from long contexts. This approach is a simpler…

eess.AS2020★ 1 cited

Mixture of Speaker-type PLDAs for Children's Speech Diarization

Jiamin Xie, Suzanna Sia, Paola Garcia +2

In diarization, the PLDA is typically used to model an inference structure which assumes the variation in speech segments be induced by various speakers. The speaker variation is t…

eess.AS2020★ 1 cited

PyChain: A Fully Parallelized PyTorch Implementation of LF-MMI for End-to-End ASR

Yiwen Shao, Yiming Wang, Daniel Povey +1

We present PyChain, a fully parallelized PyTorch implementation of end-to-end lattice-free maximum mutual information (LF-MMI) training for the so-called \emph{chain models} in the…

eess.AS2020

Multistream CNN for Robust Acoustic Modeling

Kyu J. Han, Jing Pan, Venkata Krishna Naveen Tadala +2

This paper proposes multistream CNN, a novel neural network architecture for robust acoustic modeling in speech recognition tasks. The proposed architecture processes input speech…