activity
20172022
most citedLhotse: a speech data representation library for the modern deep learning ecosystem

10 citations · 34 across the 16 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL20212 cited

speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation Assessment

Junbo Zhang, Zhiwen Zhang, Yongqing Wang +6

This paper introduces a new open-source speech corpus named "speechocean762" designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native sp…

cs.CL20213 cited

Wake Word Detection with Streaming Transformers

Yiming Wang, Hang Lv, Daniel Povey +2

Modern wake word detection systems usually rely on neural networks for acoustic modeling. Transformers has recently shown superior performance over LSTM and convolutional networks…

cs.CL2020

Efficient MDI Adaptation for n-gram Language Models

Ruizhe Huang, Ke Li, Ashish Arora +2

This paper presents an efficient algorithm for n-gram language model adaptation under the minimum discrimination information (MDI) principle, where an out-of-domain language model…

cs.CL2018

A GPU-based WFST Decoder with Exact Lattice Generation

Zhehuai Chen, Justin Luitjens, Hainan Xu +3

We describe initial work on an extension of the Kaldi toolkit that supports weighted finite-state transducer (WFST) decoding on Graphics Processing Units (GPUs). We implement token…

cs.CL20175 cited

Acoustic data-driven lexicon learning based on a greedy pronunciation selection framework

Xiaohui Zhang, Vimal Manohar, Daniel Povey +1

Speech recognition systems for irregularly-spelled languages like English normally require hand-written pronunciations. In this paper, we describe a system for automatically obtain…