4 papers
Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models
Yunkyu Lim, Jihwan Park, Hyung Yong Kim +2
Recently, Transformer-based encoder-decoder models have demonstrated strong performance in multilingual speech recognition. However, the decoder's autoregressive nature and large s…
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
Ji-Hoon Kim, Hong-Sun Yang, Yoon-Cheol Ju +3
The goal of this work is to generate natural speech in multiple languages while maintaining the same speaker identity, a task known as cross-lingual speech synthesis. A key challen…
Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting
Youkyum Kim, Jaemin Jung, Jihwan Park +2
This paper proposes a novel user-defined keyword spotting framework that accurately detects audio keywords based on text enrollment. Since audio data possesses additional acoustic…
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
Youngjoon Jang, Ji-Hoon Kim, Junseok Ahn +6
The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Spe…