10 papers
From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
Thai-Binh Nguyen, Zhaolin Li, Jan Niehues +1
Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby co…
MUSCAT: MUltilingual, SCientific ConversATion Benchmark
Supriti Sinhamahapatra, Thai-Binh Nguyen, YiÄit OÄuz +3
The goal of multilingual speech technology is to facilitate seamless communication between individuals speaking different languages, creating the experience as though everyone were…
A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
Thai-Binh Nguyen, Katerina Zmolikova, Pingchuan Ma +3
We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a s…
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
Christian Huber, Tu Anh Dinh, Carlos Mullov +10
The challenge of low-latency speech translation has recently draw significant interest in the research community as shown by several publications and shared tasks. Therefore, it is…
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
Thai-Binh Nguyen, Thi Van Nguyen, Quoc Truong Do +1
Audio-Visual Speech Recognition (AVSR) has gained significant attention recently due to its robustness against noise, which often challenges conventional speech recognition systems…
Cocktail-Party Audio-Visual Speech Recognition
Thai-Binh Nguyen, Ngoc-Quan Pham, Alexander Waibel
Audio-Visual Speech Recognition (AVSR) offers a robust solution for speech recognition in challenging environments, such as cocktail-party scenarios, where relying solely on audio…