13 papers
How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs
Christian Huber, Alexander Waibel
Recognizing new and rare words - named entities, acronyms, domain specific special words, and other items scarce in training data - remains a key challenge for automatic speech rec…
The Role of Disfluencies in Speech Translation
Maike Züfle, Maria Teleki, Fabian Retkowski +5
Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts rather than translate them.…
Adapting Foundation ASR Models to Dysarthric Speech: A Case Study
Christian Huber, Laura Kernahan, Alexander Waibel
Automatic speech recognition (ASR) systems often perform poorly in dysarthric speech, limiting their usefulness to affected speakers in everyday communication. This paper presents…
Multilingual Long-Form Speech Instruction Following: KIT's Submission to IWSLT 2026
Enes Yavuz Ugan, Maike Züfle, Yuka Ko +5
With the advent of Large Language Models, single-task and token-based multi-task models have evolved into instruction-based systems that infer task and target language implicitly f…
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
Fabian Retkowski, Maike Züfle, Thai Binh Nguyen +2
Audio chaptering, the task of segmenting long-form audio into coherent sections, is increasingly important for navigating podcasts, lectures, and videos. Despite its relevance, res…
Do What I Say: A Spoken Prompt Dataset for Instruction-Following
Maike Züfle, Sara Papi, Fabian Retkowski +5
Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompts, which may not reflect real-wo…