4 papers
Confidence-Guided Error Correction for Disordered Speech Recognition
Abner Hernandez, Tomás Arias Vergara, Andreas Maier +1
We investigate the use of large language models (LLMs) as post-processing modules for automatic speech recognition (ASR), focusing on their ability to perform error correction for…
SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis
Lukas Buess, Jan Geier, David Bani-Harouni +6
Spoken communication plays a central role in clinical workflows. In radiology, for example, most reports are created through dictation. Yet, nearly all medical AI systems rely excl…
Water Demand Forecasting of District Metered Areas through Learned Consumer Representations
Adithya Ramachandran, Thorkil Flensmark B. Neergaard, Tomás Arias-Vergara +2
Advancements in smart metering technologies have significantly improved the ability to monitor and manage water utilities. In the context of increasing uncertainty due to climate c…
A Speech-to-Video Synthesis Approach Using Spatio-Temporal Diffusion for Vocal Tract MRI
Paula Andrea Pérez-Toro, Tomás Arias-Vergara, Fangxu Xing +9
Understanding the relationship between vocal tract motion during speech and the resulting acoustic signal is crucial for aided clinical assessment and developing personalized treat…