1 citations · 2 across the 3 of their papers we have counts for
Showing eess.ASShow all
3 papers · 1 filter
eess.AS2024
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
Ruizhe Huang, Xiaohui Zhang, Zhaoheng Ni +9
Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can c…
eess.AS2023★ 1 cited
TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch
Jeff Hwang, Moto Hira, Caroline Chen +21
TorchAudio is an open-source audio and speech processing library built for PyTorch. It aims to accelerate the research and development of audio and speech technologies by providing…
eess.AS2023★ 1 cited
Exploring Speech Enhancement for Low-resource Speech Synthesis
Zhaoheng Ni, Sravya Popuri, Ning Dong +6
High-quality and intelligible speech is essential to text-to-speech (TTS) model training, however, obtaining high-quality data for low-resource languages is challenging and expensi…