2 papers
eess.AS2024
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
Hyeonseung Lee, Ji Won Yoon, Sungsoo Kim +1
Transducer neural networks have emerged as the mainstream approach for streaming automatic speech recognition (ASR), offering state-of-the-art performance in balancing accuracy and…
eess.AS2024
MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
Semin Kim, Myeonghun Jeong, Hyeonseung Lee +3
In this paper, we propose MakeSinger, a semi-supervised training method for singing voice synthesis (SVS) via classifier-free diffusion guidance. The challenge in SVS lies in the c…