2 papers
eess.AS2024
SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
Jiachen Lian, Xuanru Zhou, Zoe Ezzes +6
Speech is a hierarchical collection of text, prosody, emotions, dysfluencies, etc. Automatic transcription of speech that goes beyond text (words) is an underexplored problem. We f…
eess.AS2024
YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
Xuanru Zhou, Anshul Kashyap, Steve Li +9
Dysfluent speech detection is the bottleneck for disordered speech analysis and spoken language learning. Current state-of-the-art models are governed by rule-based systems which l…