2 papers
cs.CL2026
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis
Xi Wang, Jie Wang, Xingchen Song +8
While generative text-to-speech (TTS) models approach human-level quality, monolithic metrics fail to diagnose fine-grained acoustic artifacts or explain perceptual collapse. To ad…
cs.SD2026
Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation
Shengfan Shen, Di Wu, Xingchen Song +5
Reliable evaluation of modern zero-shot text-to-speech (TTS) models remains challenging. Subjective tests are costly and hard to reproduce, while objective metrics often saturate,…