2 papers
eess.AS2024
Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
Takaaki Saeki, Gary Wang, Nobuyuki Morioka +8
Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a…
eess.AS2023
Improving Speech Recognition for African American English With Audio Classification
Shefali Garg, Zhouyuan Huo, Khe Chai Sim +11
Automatic speech recognition (ASR) systems have been shown to have large quality disparities between the language varieties they are intended or expected to recognize. One way to m…