2 papers
eess.AS2025
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
Akam Rahimi, Triantafyllos Afouras, Andrew Zisserman
The goal of this paper is speech separation and enhancement in multi-speaker and noisy environments using a combination of different modalities. Previous works have shown good perf…
eess.AS2025
VoiceVector: Multimodal Enrolment Vectors for Speaker Separation
Akam Rahimi, Triantafyllos Afouras, Andrew Zisserman
We present a transformer-based architecture for voice separation of a target speaker from multiple other speakers and ambient noise. We achieve this by using two separate neural ne…