22 citations · 33 across the 33 of their papers we have counts for
9 papers
PauseSpeech: Natural Speech Synthesis via Pre-trained Language Model and Pause-based Prosody Modeling
Ji-Sang Hwang, Sang-Hoon Lee, Seong-Whan Lee
Although text-to-speech (TTS) systems have significantly improved, most TTS systems still have limitations in synthesizing speech with appropriate phrasing. For natural speech synt…
HiddenSinger: High-Quality Singing Voice Synthesis via Neural Audio Codec and Latent Diffusion Models
Ji-Sang Hwang, Sang-Hoon Lee, Seong-Whan Lee
Recently, denoising diffusion models have demonstrated remarkable performance among generative models in various domains. However, in the speech domain, the application of diffusio…
DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice Conversion
Ha-Yeong Choi, Sang-Hoon Lee, Seong-Whan Lee
Diffusion-based generative models have exhibited powerful generative performance in recent years. However, as many attributes exist in the data distribution and owing to several li…
OTPose: Occlusion-Aware Transformer for Pose Estimation in Sparsely-Labeled Videos
Kyung-Min Jin, Gun-Hee Lee, Seong-Whan Lee
Although many approaches for multi-human pose estimation in videos have shown profound results, they require densely annotated data which entails excessive man labor. Furthermore,…
HTNet: Anchor-free Temporal Action Localization with Hierarchical Transformers
Tae-Kyung Kang, Gun-Hee Lee, Seong-Whan Lee
Temporal action localization (TAL) is a task of identifying a set of actions in a video, which involves localizing the start and end frames and classifying each action instance. Ex…
Decoding Continual Muscle Movements Related to Complex Hand Grasping from EEG Signals
Jeong-Hyun Cho, Byoung-Hee Kwon, Byeong-Hoo Lee +1
Brain-computer interface (BCI) is a practical pathway to interpret users' intentions by decoding motor execution (ME) or motor imagery (MI) from electroencephalogram (EEG) signals.…