7 papers · 1 filter
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
Speech-to-SOAP: End-to-End Summarization of Medical Dialogues: KIT@BeTraC 2026
Enes Yavuz Ugan, Fabian Retkowski, Yuka Ko +4
With the advent of Large Language Models and its instruction following capabilities a promising application is the task of summarization. Within this domain of task the extractive…
The Role of Disfluencies in Speech Translation
Maike Züfle, Maria Teleki, Fabian Retkowski +5
Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts rather than translate them.…
Automatic Labelling of Speech Translation Errors
Dominik Macháček, Maike Züfle, Ondrej Klejch
Errors in speech translations reduce trustworthiness of Speech Translation (ST) systems and can have serious consequences. Yet currently there is no established methodology for eva…
Multilingual Long-Form Speech Instruction Following: KIT's Submission to IWSLT 2026
Enes Yavuz Ugan, Maike Züfle, Yuka Ko +5
With the advent of Large Language Models, single-task and token-based multi-task models have evolved into instruction-based systems that infer task and target language implicitly f…
When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR
Maike Züfle, Jan Niehues
SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts with sensitive information, fine-tune…