activity
20192024
most citedPhase-aware Speech Enhancement with Deep Complex U-Net

79 citations · 126 across the 12 of their papers we have counts for

collaborators

16 papers

eess.AS2024

DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance

Jinhyeok Yang, Junhyeok Lee, Hyeong-Seok Choi +3

Text-to-Speech (TTS) models have advanced significantly, aiming to accurately replicate human speech's diversity, including unique speaker identities and linguistic nuances. Despit…

cs.SD2023

Yet Another Generative Model For Room Impulse Response Estimation

Sungho Lee, Hyeong-Seok Choi, Kyogu Lee

Recent neural room impulse response (RIR) estimators typically comprise an encoder for reference audio analysis and a generator for RIR synthesis. Especially, it is the performance…

cs.SD2022

Towards trustworthy phoneme boundary detection with autoregressive model and improved evaluation metric

Hyeongju Kim, Hyeong-Seok Choi

Phoneme boundary detection has been studied due to its central role in various speech applications. In this work, we point out that this task needs to be addressed not only by algo…

cs.SD2022★ 8 cited

NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis

Hyeong-Seok Choi, Jinhyeok Yang, Juheon Lee +1

Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, most of the voice synthesis…

cs.SD2022★ 4 cited

Expressive Singing Synthesis Using Local Style Token and Dual-path Pitch Encoder

Juheon Lee, Hyeong-Seok Choi, Kyogu Lee

This paper proposes a controllable singing voice synthesis system capable of generating expressive singing voice with two novel methodologies. First, a local style token module, wh…

cs.SD2021★ 13 cited

Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations

Hyeong-Seok Choi, Juheon Lee, Wansoo Kim +3

We present a neural analysis and synthesis (NANSY) framework that can manipulate voice, pitch, and speed of an arbitrary speech signal. Most of the previous works have focused on u…