activity
20182022
most citedLhotse: a speech data representation library for the modern deep learning ecosystem

10 citations · 15 across the 12 of their papers we have counts for

collaborators

26 papers

eess.AS2022

Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation

Liyong Guo, Xiaoyu Yang, Quandong Wang +9

Knowledge distillation(KD) is a common approach to improve model performance in automatic speech recognition (ASR), where a student model is trained to imitate the output behaviour…

eess.AS20221 cited

Delay-penalized transducer for low-latency streaming ASR

Wei Kang, Zengwei Yao, Fangjun Kuang +5

In streaming automatic speech recognition (ASR), it is desirable to reduce latency as much as possible while having minimum impact on recognition accuracy. Although a few existing…

eess.AS2022

Fast and parallel decoding for transducer

Wei Kang, Liyong Guo, Fangjun Kuang +6

The transducer architecture is becoming increasingly popular in the field of speech recognition, because it is naturally streaming as well as high in accuracy. One of the drawbacks…

eess.AS20223 cited

Defense against Adversarial Attacks on Hybrid Speech Recognition using Joint Adversarial Fine-tuning with Denoiser

Sonal Joshi, Saurabh Kataria, Yiwen Shao +4

Adversarial attacks are a threat to automatic speech recognition (ASR) systems, and it becomes imperative to propose defenses to protect them. In this paper, we perform experiments…

cs.SD2022

Discovering Phonetic Inventories with Crosslingual Automatic Speech Recognition

Piotr Żelasko, Siyuan Feng, Laureano Moro Velazquez +5

The high cost of data acquisition makes Automatic Speech Recognition (ASR) model training problematic for most existing languages, including languages that do not even have a writt…

cs.SD202110 cited

Lhotse: a speech data representation library for the modern deep learning ecosystem

Piotr Żelasko, Daniel Povey, Jan "Yenda" Trmal +1

Speech data is notoriously difficult to work with due to a variety of codecs, lengths of recordings, and meta-data formats. We present Lhotse, a speech data representation library…