Showing eess.ASShow all
3 papers · 1 filter
eess.AS2025
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
Shivam Mehta, Yingru Liu, Zhenyu Tang +6
Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker…
eess.AS2024
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
Ruizhe Huang, Xiaohui Zhang, Zhaoheng Ni +9
Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can c…
eess.AS2024
Acoustic modeling for Overlapping Speech Recognition: JHU Chime-5 Challenge System
Vimal Manohar, Szu-Jui Chen, Zhiqi Wang +3
This paper summarizes our acoustic modeling efforts in the Johns Hopkins University speech recognition system for the CHiME-5 challenge to recognize highly-overlapped dinner party…