2 papers
cs.LG2025
Prompting Whisper for Improved Verbatim Transcription and End-to-end Miscue Detection
Griffin Dietz Smith, Dianna Yee, Jennifer King Chen +1
Identifying mistakes (i.e., miscues) made while reading aloud is commonly approached post-hoc by comparing automatic speech recognition (ASR) transcriptions to the target reading t…
cs.SD2025
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
Jaya Narain, Vasudha Kowtha, Colin Lea +8
Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice…