97 citations · 97 across the 3 of their papers we have counts for
Showing eess.ASShow all
2 papers · 1 filter
eess.AS2024
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
Ruizhe Huang, Xiaohui Zhang, Zhaoheng Ni +9
Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can c…
eess.AS2023
Self-Supervised Representations for Singing Voice Conversion
Tejas Jayashankar, Jilong Wu, Leda Sari +3
A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio r…