5 papers
WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models
Hanna Lee, Tan Dat Nguyen, Jaehoon Kang +1
Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale quadratically with sequence length due to f…
GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech
Jaehoon Kang, Yejin Lee, Kyuhong Shim
We propose GLASS, a framework for composable acoustic style control in zero-shot autoregressive text-to-speech (TTS) that learns controls from post-generation rewards rather than s…
Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models
Jaehoon Kang, Yejin Lee, Yoonji Park +1
While prompt-based text-to-speech (TTS) models enable natural language-driven speaking style control, they often provide limited fine-grained control and apply a single global styl…
DialBGM: A Benchmark for Background Music Recommendation from Everyday Multi-Turn Dialogues
Joonhyeok Shin, Jaehoon Kang, Yujun Lee +4
Selecting an appropriate background music (BGM) that supports natural human conversation is a common production step in media and interactive systems. In this paper, we introduce d…
P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-Speech
Yejin Lee, Jaehoon Kang, Kyuhong Shim
While persona-driven large language models (LLMs) and prompt-based text-to-speech (TTS) systems have advanced significantly, a usability gap arises when users attempt to generate v…