79 citations · 126 across the 12 of their papers we have counts for
16 papers
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
Jinhyeok Yang, Junhyeok Lee, Hyeong-Seok Choi +3
Text-to-Speech (TTS) models have advanced significantly, aiming to accurately replicate human speech's diversity, including unique speaker identities and linguistic nuances. Despit…
Yet Another Generative Model For Room Impulse Response Estimation
Sungho Lee, Hyeong-Seok Choi, Kyogu Lee
Recent neural room impulse response (RIR) estimators typically comprise an encoder for reference audio analysis and a generator for RIR synthesis. Especially, it is the performance…
Towards trustworthy phoneme boundary detection with autoregressive model and improved evaluation metric
Hyeongju Kim, Hyeong-Seok Choi
Phoneme boundary detection has been studied due to its central role in various speech applications. In this work, we point out that this task needs to be addressed not only by algo…
NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis
Hyeong-Seok Choi, Jinhyeok Yang, Juheon Lee +1
Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, most of the voice synthesis…
Expressive Singing Synthesis Using Local Style Token and Dual-path Pitch Encoder
Juheon Lee, Hyeong-Seok Choi, Kyogu Lee
This paper proposes a controllable singing voice synthesis system capable of generating expressive singing voice with two novel methodologies. First, a local style token module, wh…
Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations
Hyeong-Seok Choi, Juheon Lee, Wansoo Kim +3
We present a neural analysis and synthesis (NANSY) framework that can manipulate voice, pitch, and speed of an arbitrary speech signal. Most of the previous works have focused on u…