activity
20132022
most citedThe Hitachi-JHU DIHARD III System: Competitive End-to-End Neural Diarization and X-Vector Clustering Systems Combined by DOVER-Lap

27 citations · 83 across the 26 of their papers we have counts for

collaborators

35 papers

eess.AS20223 cited

Defense against Adversarial Attacks on Hybrid Speech Recognition using Joint Adversarial Fine-tuning with Denoiser

Sonal Joshi, Saurabh Kataria, Yiwen Shao +4

Adversarial attacks are a threat to automatic speech recognition (ASR) systems, and it becomes imperative to propose defenses to protect them. In this paper, we perform experiments…

eess.AS2022

PHO-LID: A Unified Model Incorporating Acoustic-Phonetic and Phonotactic Information for Language Identification

Hexin Liu, Leibny Paola Garcia Perera, Andy W. H. Khong +2

We propose a novel model to hierarchically incorporate phoneme and phonotactic information for language identification (LID) without requiring phoneme annotations for training. In…

eess.AS2022

Investigating self-supervised learning for speech enhancement and separation

Zili Huang, Shinji Watanabe, Shu-wen Yang +2

Speech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target spe…

eess.AS20221 cited

Enhance Language Identification using Dual-mode Model with Knowledge Distillation

Hexin Liu, Leibny Paola Garcia Perera, Andy W. H. Khong +3

In this paper, we propose to employ a dual-mode framework on the x-vector self-attention (XSA-LID) model with knowledge distillation (KD) to enhance its language identification (LI…

cs.SD202110 cited

Lhotse: a speech data representation library for the modern deep learning ecosystem

Piotr Żelasko, Daniel Povey, Jan "Yenda" Trmal +1

Speech data is notoriously difficult to work with due to a variety of codecs, lengths of recordings, and meta-data formats. We present Lhotse, a speech data representation library…

eess.AS2021

Injecting Text and Cross-lingual Supervision in Few-shot Learning from Self-Supervised Models

Matthew Wiesner, Desh Raj, Sanjeev Khudanpur

Self-supervised model pre-training has recently garnered significant interest, but relatively few efforts have explored using additional resources in fine-tuning these models. We d…