7 papers · 1 filter
From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
Thai-Binh Nguyen, Zhaolin Li, Jan Niehues +1
Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby co…
Multimodal In-context Learning for ASR of Low-resource Languages
Zhaolin Li, Jan Niehues
Automatic speech recognition (ASR) still covers only a small fraction of the world's languages, mainly due to supervised data scarcity. In-context learning (ICL) with large languag…
In-context Language Learning for Endangered Languages in Speech Recognition
Zhaolin Li, Jan Niehues
With approximately 7,000 languages spoken worldwide, current large language models (LLMs) support only a small subset. Prior research indicates LLMs can learn new languages for cer…
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
Zhaolin Li, Yining Liu, Danni Liu +6
This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translati…
Augmenting Automatic Speech Recognition Models with Disfluency Detection
Robin Amann, Zhaolin Li, Barbara Bruno +1
Speech disfluency commonly occurs in conversational and spontaneous speech. However, standard Automatic Speech Recognition (ASR) models struggle to accurately recognize these disfl…
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
Sai Koneru, Thai-Binh Nguyen, Ngoc-Quan Pham +4
Large Language Models (LLMs) are currently under exploration for various tasks, including Automatic Speech Recognition (ASR), Machine Translation (MT), and even End-to-End Speech T…