activity
20242026
collaborators

5 papers

cs.CL2026

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

Thai-Binh Nguyen, Zhaolin Li, Jan Niehues +1

Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby co…

cs.CL2026

Multimodal In-context Learning for ASR of Low-resource Languages

Zhaolin Li, Jan Niehues

Automatic speech recognition (ASR) still covers only a small fraction of the world's languages, mainly due to supervised data scarcity. In-context learning (ICL) with large languag…

cs.CL2026

In-context Language Learning for Endangered Languages in Speech Recognition

Zhaolin Li, Jan Niehues

With approximately 7,000 languages spoken worldwide, current large language models (LLMs) support only a small subset. Prior research indicates LLMs can learn new languages for cer…

cs.CL2025

KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization

Zhaolin Li, Yining Liu, Danni Liu +6

This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translati…

cs.CL2024

SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading

Tu Anh Dinh, Carlos Mullov, Leonard Bärmann +15

With the rapid development of Large Language Models (LLMs), it is crucial to have benchmarks which can evaluate the ability of LLMs on different domains. One common use of LLMs is…