activity
20182022
most citedLhotse: a speech data representation library for the modern deep learning ecosystem

10 citations · 15 across the 12 of their papers we have counts for

collaborators
Showing eess.ASShow all

15 papers · 1 filter

eess.AS2022

Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation

Liyong Guo, Xiaoyu Yang, Quandong Wang +9

Knowledge distillation(KD) is a common approach to improve model performance in automatic speech recognition (ASR), where a student model is trained to imitate the output behaviour…

eess.AS20221 cited

Delay-penalized transducer for low-latency streaming ASR

Wei Kang, Zengwei Yao, Fangjun Kuang +5

In streaming automatic speech recognition (ASR), it is desirable to reduce latency as much as possible while having minimum impact on recognition accuracy. Although a few existing…

eess.AS2022

Fast and parallel decoding for transducer

Wei Kang, Liyong Guo, Fangjun Kuang +6

The transducer architecture is becoming increasingly popular in the field of speech recognition, because it is naturally streaming as well as high in accuracy. One of the drawbacks…

eess.AS20223 cited

Defense against Adversarial Attacks on Hybrid Speech Recognition using Joint Adversarial Fine-tuning with Denoiser

Sonal Joshi, Saurabh Kataria, Yiwen Shao +4

Adversarial attacks are a threat to automatic speech recognition (ASR) systems, and it becomes imperative to propose defenses to protect them. In this paper, we perform experiments…

eess.AS2021

Unsupervised Speech Segmentation and Variable Rate Representation Learning using Segmental Contrastive Predictive Coding

Saurabhchand Bhati, Jesús Villalba, Piotr Żelasko +2

Typically, unsupervised segmentation of speech into the phone and word-like units are treated as separate tasks and are often done via different methods which do not fully leverage…

eess.AS2021

Representation Learning to Classify and Detect Adversarial Attacks against Speaker and Speech Recognition Systems

Jesús Villalba, Sonal Joshi, Piotr Żelasko +1

Adversarial attacks have become a major threat for machine learning applications. There is a growing interest in studying these attacks in the audio domain, e.g, speech and speaker…