activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

Thai-Binh Nguyen, Zhaolin Li, Jan Niehues +1

Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby co…

cs.CL2026

Multimodal In-context Learning for ASR of Low-resource Languages

Zhaolin Li, Jan Niehues

Automatic speech recognition (ASR) still covers only a small fraction of the world's languages, mainly due to supervised data scarcity. In-context learning (ICL) with large languag…

cs.CL2025

In-context Language Learning for Endangered Languages in Speech Recognition

Zhaolin Li, Jan Niehues

With approximately 7,000 languages spoken worldwide, current large language models (LLMs) support only a small subset. Prior research indicates LLMs can learn new languages for cer…

cs.CL2025

KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization

Zhaolin Li, Yining Liu, Danni Liu +6

This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translati…

cs.CL2024

Augmenting Automatic Speech Recognition Models with Disfluency Detection

Robin Amann, Zhaolin Li, Barbara Bruno +1

Speech disfluency commonly occurs in conversational and spontaneous speech. However, standard Automatic Speech Recognition (ASR) models struggle to accurately recognize these disfl…

cs.CL2024

Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024

Sai Koneru, Thai-Binh Nguyen, Ngoc-Quan Pham +4

Large Language Models (LLMs) are currently under exploration for various tasks, including Automatic Speech Recognition (ASR), Machine Translation (MT), and even End-to-End Speech T…