paper

Benchmarking Commercial Speech Recognition and Multimodal Large Language Models on Dysarthric Speech: Severity-Stratified Baselines and Architecture-Specific Prompting Effects

arXiv:2512.17474 · doi:10.1155/int/6065038

Abstract

Voice-based human-machine interaction has become a primary means of accessing intelligent systems, yet individuals with dysarthria are systematically excluded by persistent gaps in recognition accuracy. Although automatic speech recognition (ASR) achieves word error rates (WER) below 5% on typical speech, performance degrades sharply for dysarthric speakers, while the zero-shot behaviour of multimodal large language models (MLLMs) on such speech remains unclear. We evaluate eight commercial speech-to-text services on the TORGO dysarthric speech corpus: four conventional ASR systems (AssemblyAI, Whisper large-v3, Deepgram Nova-3, Nova-3 Medical) and four MLLM-based systems (GPT-4o, GPT-4o Mini, Gemini 2.5 Pro, Gemini 2.5 Flash), using lexical accuracy, semantic preservation, and cost-latency measures. Recognition degraded consistently with severity. Mild dysarthria reached low single-digit WER, around 1-2% for the leading systems, whereas severe dysarthria exceeded 51% WER for every system, with no MLLM advantage over conventional ASR under default settings. A four-condition prompt ablation showed architecture-specific effects: for the OpenAI models, verbatim-transcription prompts reduced severe-tier WER mainly by suppressing non-target-language drift, lowering GPT-4o from 60.1% to 52.9% and GPT-4o Mini from 66.0% to about 55%; Gemini models showed no consistent benefit and sometimes degraded. Semantic metrics correlated strongly with WER and were largely redundant in aggregate, but identified cases where communicative intent was partly preserved despite poor lexical accuracy. These severity-stratified, per-speaker baselines provide a reusable reference for evidence-based technology selection in assistive voice interfaces.

32 pages. Revised peer-reviewed version. Updated dataset filtering and transcript normalization, expanded four-condition prompting ablation and error analysis, and revised statistical reporting. Published in International Journal of Intelligent Systems

Benchmarking Commercial Speech Recognition and Multimodal Large Language Models on Dysarthric Speech: Severity-Stratified Baselines and Architecture-Specific Prompting Effects · wovepaper