15 papers
From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
Thai-Binh Nguyen, Zhaolin Li, Jan Niehues +1
Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby co…
Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS
Seymanur Akti, Alexander Waibel
Humans tend to speak louder and clearer in challenging environments, such as noisy conditions or when addressing hearingimpaired listeners, which is called Lombard effect. To simul…
Adding Robust Code-Switching Capabilities to High Performance Multilingual ASR
Enes Yavuz Ugan, Alexander Waibel
Code-switching (CSW) remains challenging for large multi-lingual ASR systems in real-world deployment. While fine-tuning on synthetic CSW data is possible, it generally degrades st…
KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026
Seymanur Akti, Alexander Waibel
Cross-lingual voice cloning aims to generate speech in a target language while preserving speaker identity from a source-language reference. This task is central to speech translat…
A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
Thai-Binh Nguyen, Katerina Zmolikova, Pingchuan Ma +3
We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a s…
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
Seymanur Akti, Alexander Waibel
The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-sp…