activity
20172022
most citedLhotse: a speech data representation library for the modern deep learning ecosystem

10 citations · 34 across the 16 of their papers we have counts for

collaborators
Showing eess.ASShow all

17 papers · 1 filter

eess.AS2024

Less Peaky and More Accurate CTC Forced Alignment by Label Priors

Ruizhe Huang, Xiaohui Zhang, Zhaoheng Ni +9

Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can c…

eess.AS2024

On Speaker Attribution with SURT

Desh Raj, Matthew Wiesner, Matthew Maciejewski +3

The Streaming Unmixing and Recognition Transducer (SURT) has recently become a popular framework for continuous, streaming, multi-talker speech recognition (ASR). With advances in…

eess.AS20231 cited

Learning from Flawed Data: Weakly Supervised Automatic Speech Recognition

Dongji Gao, Hainan Xu, Desh Raj +3

Training automatic speech recognition (ASR) systems requires large amounts of well-curated paired data. However, human annotators usually perform "non-verbatim" transcription, whic…

eess.AS20232 cited

Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition

Han Zhu, Dongji Gao, Gaofeng Cheng +3

When labeled data is insufficient, semi-supervised learning with the pseudo-labeling technique can significantly improve the performance of automatic speech recognition. However, p…

eess.AS2022

Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation

Liyong Guo, Xiaoyu Yang, Quandong Wang +9

Knowledge distillation(KD) is a common approach to improve model performance in automatic speech recognition (ASR), where a student model is trained to imitate the output behaviour…

eess.AS20221 cited

Delay-penalized transducer for low-latency streaming ASR

Wei Kang, Zengwei Yao, Fangjun Kuang +5

In streaming automatic speech recognition (ASR), it is desirable to reduce latency as much as possible while having minimum impact on recognition accuracy. Although a few existing…