activity
20132026
most citedThe Hitachi-JHU DIHARD III System: Competitive End-to-End Neural Diarization and X-Vector Clustering Systems Combined by DOVER-Lap

27 citations · 104 across the 68 of their papers we have counts for

collaborators
Showing 2022 · eess.ASShow all

7 papers · 2 filters

eess.AS2022★ 5 cited

GPU-accelerated Guided Source Separation for Meeting Transcription

Desh Raj, Daniel Povey, Sanjeev Khudanpur

Guided source separation (GSS) is a type of target-speaker extraction method that relies on pre-computed speaker activities and blind source separation to perform front-end enhance…

eess.AS2022★ 2 cited

Adapting self-supervised models to multi-talker speech recognition using speaker embeddings

Zili Huang, Desh Raj, Paola García +1

Self-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-t…

eess.AS2022★ 1 cited

Reducing Language confusion for Code-switching Speech Recognition with Token-level Language Diarization

Hexin Liu, Haihua Xu, Leibny Paola Garcia +3

Code-switching (CS) refers to the phenomenon that languages switch within a speech signal and leads to language confusion for automatic speech recognition (ASR). This paper aims to…

eess.AS2022★ 3 cited

Defense against Adversarial Attacks on Hybrid Speech Recognition using Joint Adversarial Fine-tuning with Denoiser

Sonal Joshi, Saurabh Kataria, Yiwen Shao +4

Adversarial attacks are a threat to automatic speech recognition (ASR) systems, and it becomes imperative to propose defenses to protect them. In this paper, we perform experiments…

eess.AS2022

PHO-LID: A Unified Model Incorporating Acoustic-Phonetic and Phonotactic Information for Language Identification

Hexin Liu, Leibny Paola Garcia Perera, Andy W. H. Khong +2

We propose a novel model to hierarchically incorporate phoneme and phonotactic information for language identification (LID) without requiring phoneme annotations for training. In…

eess.AS2022

Investigating self-supervised learning for speech enhancement and separation

Zili Huang, Shinji Watanabe, Shu-wen Yang +2

Speech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target spe…