1 paper
Xiao Zhou, Oisín Turbitt, Kit Bower-Morris +3
Decoder-only text-to-speech (TTS) models scale efficiently but remain prone to content hallucinations that arise from weak text-speech alignment during autoregressive generation. W…