4 papers
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
Bruno Korbar, Andrew Zisserman
The goal of this paper is to be able to retrieve images using a compound query that combines object instance information from an image, with a natural text description of what that…
Understanding Co-speech Gestures in-the-wild
Sindhu B Hegde, K R Prajwal, Taein Kwon +1
Co-speech gestures play a vital role in non-verbal communication. In this paper, we introduce a new framework for co-speech gesture understanding in the wild. Specifically, we prop…
VoiceVector: Multimodal Enrolment Vectors for Speaker Separation
Akam Rahimi, Triantafyllos Afouras, Andrew Zisserman
We present a transformer-based architecture for voice separation of a target speaker from multiple other speakers and ambient noise. We achieve this by using two separate neural ne…
New keypoint-based approach for recognising British Sign Language (BSL) from sequences
Oishi Deb, KR Prajwal, Andrew Zisserman
In this paper, we present a novel keypoint-based classification model designed to recognise British Sign Language (BSL) words within continuous signing sequences. Our model's perfo…