activity
20242026
collaborators

10 papers

cs.CL2026

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

Thai-Binh Nguyen, Zhaolin Li, Jan Niehues +1

Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby co…

cs.CL2026

MUSCAT: MUltilingual, SCientific ConversATion Benchmark

Supriti Sinhamahapatra, Thai-Binh Nguyen, Yiğit Oğuz +3

The goal of multilingual speech technology is to facilitate seamless communication between individuals speaking different languages, creating the experience as though everyone were…

cs.CL2026

A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results

Thai-Binh Nguyen, Katerina Zmolikova, Pingchuan Ma +3

We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a s…

cs.CL2025

End-to-End Evaluation for Low-Latency Simultaneous Speech Translation

Christian Huber, Tu Anh Dinh, Carlos Mullov +10

The challenge of low-latency speech translation has recently draw significant interest in the research community as shown by several publications and shared tasks. Therefore, it is…

cs.CL2025

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition

Thai-Binh Nguyen, Thi Van Nguyen, Quoc Truong Do +1

Audio-Visual Speech Recognition (AVSR) has gained significant attention recently due to its robustness against noise, which often challenges conventional speech recognition systems…

cs.SD2025

Cocktail-Party Audio-Visual Speech Recognition

Thai-Binh Nguyen, Ngoc-Quan Pham, Alexander Waibel

Audio-Visual Speech Recognition (AVSR) offers a robust solution for speech recognition in challenging environments, such as cocktail-party scenarios, where relying solely on audio…