collaborators

6 papers

cs.SD2026

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching

Jinhyeok Yang, Hyeongju Kim, Yechan Yu +3

While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repe…

cs.SD2025

Robust TTS Training via Self-Purifying Flow Matching for the WildSpoof 2026 TTS Track

June Young Yi, Hyeongju Kim, Juheon Lee

This paper presents a lightweight text-to-speech (TTS) system developed for the WildSpoof Challenge TTS Track. Our approach fine-tunes the recently released open-weight TTS model,…

eess.AS2025

Improving Test-Time Performance of RVQ-based Neural Codecs

Hyeongju Kim, Junhyeok Lee, Jacob Morton +2

The residual vector quantization (RVQ) technique plays a central role in recent advances in neural audio codecs. These models effectively synthesize high-fidelity audio from a limi…

eess.AS2025

Training Flow Matching Models with Reliable Labels via Self-Purification

Hyeongju Kim, Yechan Yu, June Young Yi +1

Training datasets are inherently imperfect, often containing mislabeled samples due to human annotation errors, limitations of tagging models, and other sources of noise. Such labe…

eess.AS2025

SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System

Hyeongju Kim, Jinhyeok Yang, Yechan Yu +5

We introduce SupertonicTTS, a novel text-to-speech (TTS) system designed for efficient and streamlined speech synthesis. SupertonicTTS comprises three components: a speech autoenco…

eess.AS2025

Length-Aware Rotary Position Embedding for Text-Speech Alignment

Hyeongju Kim, Juheon Lee, Jinhyeok Yang +1

Many recent text-to-speech (TTS) systems are built on transformer architectures and employ cross-attention mechanisms for text-speech alignment. Within these systems, rotary positi…