3 papers
cs.LG2025
Prompting Whisper for Improved Verbatim Transcription and End-to-end Miscue Detection
Griffin Dietz Smith, Dianna Yee, Jennifer King Chen +1
Identifying mistakes (i.e., miscues) made while reading aloud is commonly approached post-hoc by comparing automatic speech recognition (ASR) transcriptions to the target reading t…
cs.SD2025
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
Jaya Narain, Vasudha Kowtha, Colin Lea +8
Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice…
cs.LG2024
Hypernetworks for Personalizing ASR to Atypical Speech
Max Müller-Eberstein, Dianna Yee, Karren Yang +2
Parameter-efficient fine-tuning (PEFT) for personalizing automatic speech recognition (ASR) has recently shown promise for adapting general population models to atypical speech. Ho…