27 citations · 83 across the 26 of their papers we have counts for
35 papers
Defense against Adversarial Attacks on Hybrid Speech Recognition using Joint Adversarial Fine-tuning with Denoiser
Sonal Joshi, Saurabh Kataria, Yiwen Shao +4
Adversarial attacks are a threat to automatic speech recognition (ASR) systems, and it becomes imperative to propose defenses to protect them. In this paper, we perform experiments…
PHO-LID: A Unified Model Incorporating Acoustic-Phonetic and Phonotactic Information for Language Identification
Hexin Liu, Leibny Paola Garcia Perera, Andy W. H. Khong +2
We propose a novel model to hierarchically incorporate phoneme and phonotactic information for language identification (LID) without requiring phoneme annotations for training. In…
Investigating self-supervised learning for speech enhancement and separation
Zili Huang, Shinji Watanabe, Shu-wen Yang +2
Speech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target spe…
Enhance Language Identification using Dual-mode Model with Knowledge Distillation
Hexin Liu, Leibny Paola Garcia Perera, Andy W. H. Khong +3
In this paper, we propose to employ a dual-mode framework on the x-vector self-attention (XSA-LID) model with knowledge distillation (KD) to enhance its language identification (LI…
Lhotse: a speech data representation library for the modern deep learning ecosystem
Piotr Żelasko, Daniel Povey, Jan "Yenda" Trmal +1
Speech data is notoriously difficult to work with due to a variety of codecs, lengths of recordings, and meta-data formats. We present Lhotse, a speech data representation library…
Injecting Text and Cross-lingual Supervision in Few-shot Learning from Self-Supervised Models
Matthew Wiesner, Desh Raj, Sanjeev Khudanpur
Self-supervised model pre-training has recently garnered significant interest, but relatively few efforts have explored using additional resources in fine-tuning these models. We d…