2 papers
cs.SD2025
Shared Multi-modal Embedding Space for Face-Voice Association
Christopher Simic, Korbinian Riedhammer, Tobias Bocklet
The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model w…
cs.SD2025
Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
Christopher Simic, Korbinian Riedhammer, Tobias Bocklet
We present an approach to Audio-Visual Speech Recognition that builds on a pre-trained Whisper model. To infuse visual information into this audio-only model, we extend it with an…