8 papers
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
Jinhyeok Yang, Hyeongju Kim, Yechan Yu +3
While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repe…
Super Monotonic Alignment Search
Junhyeok Lee, Hyeongju Kim
Monotonic alignment search (MAS), introduced by Glow-TTS, is one of the most popular algorithm in text-to-speech to estimate unknown alignments between text and speech. Since this…
Robust TTS Training via Self-Purifying Flow Matching for the WildSpoof 2026 TTS Track
June Young Yi, Hyeongju Kim, Juheon Lee
This paper presents a lightweight text-to-speech (TTS) system developed for the WildSpoof Challenge TTS Track. Our approach fine-tunes the recently released open-weight TTS model,…
Improving Test-Time Performance of RVQ-based Neural Codecs
Hyeongju Kim, Junhyeok Lee, Jacob Morton +2
The residual vector quantization (RVQ) technique plays a central role in recent advances in neural audio codecs. These models effectively synthesize high-fidelity audio from a limi…
Training Flow Matching Models with Reliable Labels via Self-Purification
Hyeongju Kim, Yechan Yu, June Young Yi +1
Training datasets are inherently imperfect, often containing mislabeled samples due to human annotation errors, limitations of tagging models, and other sources of noise. Such labe…
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
Hyeongju Kim, Jinhyeok Yang, Yechan Yu +5
We introduce SupertonicTTS, a novel text-to-speech (TTS) system designed for efficient and streamlined speech synthesis. SupertonicTTS comprises three components: a speech autoenco…