2 papers
cs.SD2026
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
Jia-Hong Huang, Seulgi Kim, Yi Chieh Liu +5
Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker id…
cs.SD2026
No Word Left Behind: Mitigating Prefix Bias in Open-Vocabulary Keyword Spotting
Yi Liu, Chuan-Che Huang, Xiao Quan
Open-vocabulary keyword spotting (OV-KWS) enables personalized device control via arbitrary voice commands. Recently, researchers have explored using audio-text joint embeddings, a…