11 citations · 18 across the 7 of their papers we have counts for
5 papers · 1 filter
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
Shivam Mehta, Yingru Liu, Zhenyu Tang +6
Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker…
Accent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings
Jialu Li, Vimal Manohar, Pooja Chitkara +5
Speech recognition models often obtain degraded performance when tested on speech with unseen accents. Domain-adversarial training (DAT) and multi-task learning (MTL) are two commo…
On lattice-free boosted MMI training of HMM and CTC-based full-context ASR models
Xiaohui Zhang, Vimal Manohar, David Zhang +7
Hybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria. However, they have vastly different legacies and are usually implem…
Kaizen: Continuously improving teacher using Exponential Moving Average for semi-supervised speech recognition
Vimal Manohar, Tatiana Likhomanenko, Qiantong Xu +5
In this paper, we introduce the Kaizen framework that uses a continuously improving teacher to generate pseudo-labels for semi-supervised speech recognition (ASR). The proposed app…
Large scale weakly and semi-supervised learning for low-resource video ASR
Kritika Singh, Vimal Manohar, Alex Xiao +7
Many semi- and weakly-supervised approaches have been investigated for overcoming the labeling cost of building high quality speech recognition systems. On the challenging task of…