10 citations · 15 across the 12 of their papers we have counts for
26 papers
Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation
Liyong Guo, Xiaoyu Yang, Quandong Wang +9
Knowledge distillation(KD) is a common approach to improve model performance in automatic speech recognition (ASR), where a student model is trained to imitate the output behaviour…
Delay-penalized transducer for low-latency streaming ASR
Wei Kang, Zengwei Yao, Fangjun Kuang +5
In streaming automatic speech recognition (ASR), it is desirable to reduce latency as much as possible while having minimum impact on recognition accuracy. Although a few existing…
Fast and parallel decoding for transducer
Wei Kang, Liyong Guo, Fangjun Kuang +6
The transducer architecture is becoming increasingly popular in the field of speech recognition, because it is naturally streaming as well as high in accuracy. One of the drawbacks…
Defense against Adversarial Attacks on Hybrid Speech Recognition using Joint Adversarial Fine-tuning with Denoiser
Sonal Joshi, Saurabh Kataria, Yiwen Shao +4
Adversarial attacks are a threat to automatic speech recognition (ASR) systems, and it becomes imperative to propose defenses to protect them. In this paper, we perform experiments…
Discovering Phonetic Inventories with Crosslingual Automatic Speech Recognition
Piotr Żelasko, Siyuan Feng, Laureano Moro Velazquez +5
The high cost of data acquisition makes Automatic Speech Recognition (ASR) model training problematic for most existing languages, including languages that do not even have a writt…
Lhotse: a speech data representation library for the modern deep learning ecosystem
Piotr Żelasko, Daniel Povey, Jan "Yenda" Trmal +1
Speech data is notoriously difficult to work with due to a variety of codecs, lengths of recordings, and meta-data formats. We present Lhotse, a speech data representation library…