2.9k citations · 4.7k across the 57 of their papers we have counts for
3 papers · 1 filter
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
Akam Rahimi, Triantafyllos Afouras, Andrew Zisserman
The goal of this paper is speech separation and enhancement in multi-speaker and noisy environments using a combination of different modalities. Previous works have shown good perf…
Disentangled Speech Embeddings using Cross-modal Self-supervision
Arsha Nagrani, Joon Son Chung, Samuel Albanie +1
The objective of this paper is to learn representations of speaker identity without access to manually annotated data. To do so, we develop a self-supervised learning objective tha…
Utterance-level Aggregation For Speaker Recognition In The Wild
Weidi Xie, Arsha Nagrani, Joon Son Chung +1
The objective of this paper is speaker recognition "in the wild"-where utterances may be of variable length and also contain irrelevant signals. Crucial elements in the design of d…