27 citations · 35 across the 7 of their papers we have counts for
11 papers · 1 filter
Adapting self-supervised models to multi-talker speech recognition using speaker embeddings
Zili Huang, Desh Raj, Paola García +1
Self-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-t…
Injecting Text and Cross-lingual Supervision in Few-shot Learning from Self-Supervised Models
Matthew Wiesner, Desh Raj, Sanjeev Khudanpur
Self-supervised model pre-training has recently garnered significant interest, but relatively few efforts have explored using additional resources in fine-tuning these models. We d…
Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker
Maokui He, Desh Raj, Zili Huang +3
Target-speaker voice activity detection (TS-VAD) has recently shown promising results for speaker diarization on highly overlapped speech. However, the original model requires a fi…
Reformulating DOVER-Lap Label Mapping as a Graph Partitioning Problem
Desh Raj, Sanjeev Khudanpur
We recently proposed DOVER-Lap, a method for combining overlap-aware speaker diarization system outputs. DOVER-Lap improved upon its predecessor DOVER by using a label mapping meth…
The Hitachi-JHU DIHARD III System: Competitive End-to-End Neural Diarization and X-Vector Clustering Systems Combined by DOVER-Lap
Shota Horiguchi, Nelson Yalta, Paola Garcia +7
This paper provides a detailed description of the Hitachi-JHU system that was submitted to the Third DIHARD Speech Diarization Challenge. The system outputs the ensemble results of…
Multi-class Spectral Clustering with Overlaps for Speaker Diarization
Desh Raj, Zili Huang, Sanjeev Khudanpur
This paper describes a method for overlap-aware speaker diarization. Given an overlap detector and a speaker embedding extractor, our method performs spectral clustering of segment…