3 papers
cs.SD2026
From Reliable Text to Real Voices: Trust-Aware Progressive Adaptation for Low-Resource TTS
Jiayi Lu, Yizhong Geng, Jinghan Yang +4
Low-resource text-to-speech (TTS) adaptation is constrained by scarce paired data and costly manual transcription. Existing fixed-voice TTS systems can provide relatively accurate…
cs.SD2026
SceneTTS-Bench: A Benchmark for Scene-Level TTS in Drama Dubbing
Yizhong Geng, Yanliang Li, Jinghan Yang +3
Text-to-speech systems are increasingly used for drama dubbing, yet evaluation protocols remain sentence-level, leaving critical scene-level behaviors insufficiently measured. We p…
cs.CL2026
Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models
Yizhong Geng, Yanliang Li, Jinghan Yang +4
Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However, their effectiveness in low-res…