3 papers
cs.CL2026
Hearing Between the Lines: Unlocking the Reasoning Power of LLMs for Speech Evaluation
Arjun Chandra, Kevin Miller, Venkatesh Ravichandran +2
Large Language Model (LLM) judges exhibit strong reasoning capabilities but are limited to textual content. This leaves current automatic Speech-to-Speech (S2S) evaluation methods…
eess.AS2025
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
Xinlu He, Swayambhu Nath Ray, Harish Mallidi +5
Unified architectures in multimodal large language models (MLLM) have shown promise in handling diverse tasks within a single framework. In the text-to-speech (TTS) task, current M…
cs.AI2025
The Interspeech 2025 Speech Accessibility Project Challenge
Xiuwen Zheng, Bornali Phukon, Jonghwan Na +13
While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities re…