11 papers
Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS
Seymanur Akti, Alexander Waibel
Humans tend to speak louder and clearer in challenging environments, such as noisy conditions or when addressing hearingimpaired listeners, which is called Lombard effect. To simul…
KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026
Seymanur Akti, Alexander Waibel
Cross-lingual voice cloning aims to generate speech in a target language while preserving speaker identity from a source-language reference. This task is central to speech translat…
Multilingual Long-Form Speech Instruction Following: KIT's Submission to IWSLT 2026
Enes Yavuz Ugan, Maike Züfle, Yuka Ko +5
With the advent of Large Language Models, single-task and token-based multi-task models have evolved into instruction-based systems that infer task and target language implicitly f…
BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion
Sai Koneru, Fabian Retkowski, Christian Huber +5
The globalization of education and rapid growth of online learning have made localizing educational content a critical challenge. Lecture materials are inherently multimodal, combi…
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
Seymanur Akti, Alexander Waibel
The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-sp…
Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
Dogucan Yaman, Seymanur Akti, Fevziye Irem Eyiokur +1
We propose a text-to-talking-face synthesis framework leveraging latent speech representations from HierSpeech++. A Text-to-Vec module generates Wav2Vec2 embeddings from text, whic…