2 papers
cs.SD2026
Diagnose, Then Refine: A Closed-Loop TTS System with AudioLLM-Guided Correction
Zeyang Song, Tianchi Liu, Tianrui Wang +3
Current TTS systems typically rely on open-loop, single-pass generation and can produce sporadic local prosodic defects, such as misplaced stress, unnatural pauses, or flattened in…
cs.CL2026
SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings
Priyam Mazumdar, Yurii Halychanskyi, Steven Guo +2
Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics…