3 papers
cs.CL2026
The Anatomy of an ASR Hallucination
Hamees Sayed, Apoorv Singh, Kumar Aman +1
ASR systems sometimes produce fluent text that is unrelated to the speech they receive. We view these hallucinations as one possible consequence of a broader grounding failure, in…
cs.CL2026
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
Wasim Madha, Nityanand Mathur, Hamees Sayed +4
Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and deco…
cs.AI2026
How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech
Nityanand Mathur, Hamees Sayed, Wasim Madha +4
Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this…