6 papers
Recognising BSL Fingerspelling in Continuous Signing Sequences
Alyssa Chan, Taein Kwon, Andrew Zisserman
Fingerspelling is a critical component of British Sign Language (BSL), used to spell proper names, technical terms, and words that lack established lexical signs. Fingerspelling re…
Scaling Audio-Text Retrieval with Multimodal Large Language Models
Jilan Xu, Carl Thomé, Danijela Horak +2
Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentall…
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
Bruno Korbar, Andrew Zisserman
The goal of this paper is to be able to retrieve images using a compound query that combines object instance information from an image, with a natural text description of what that…
Understanding Co-speech Gestures in-the-wild
Sindhu B Hegde, K R Prajwal, Taein Kwon +1
Co-speech gestures play a vital role in non-verbal communication. In this paper, we introduce a new framework for co-speech gesture understanding in the wild. Specifically, we prop…
VoiceVector: Multimodal Enrolment Vectors for Speaker Separation
Akam Rahimi, Triantafyllos Afouras, Andrew Zisserman
We present a transformer-based architecture for voice separation of a target speaker from multiple other speakers and ambient noise. We achieve this by using two separate neural ne…
New keypoint-based approach for recognising British Sign Language (BSL) from sequences
Oishi Deb, KR Prajwal, Andrew Zisserman
In this paper, we present a novel keypoint-based classification model designed to recognise British Sign Language (BSL) words within continuous signing sequences. Our model's perfo…