1 paper
Chenlin Liu, Minghui Fang, Zhonghao Bi +3
Language model-based text-to-speech (LM-based TTS) remains vulnerable to speech hallucinations that deviate from the target text. Existing mitigation mainly relies on architectural…