activity
20242026
collaborators

6 papers

cs.CV2026

Recognising BSL Fingerspelling in Continuous Signing Sequences

Alyssa Chan, Taein Kwon, Andrew Zisserman

Fingerspelling is a critical component of British Sign Language (BSL), used to spell proper names, technical terms, and words that lack established lexical signs. Fingerspelling re…

cs.SD2026

Scaling Audio-Text Retrieval with Multimodal Large Language Models

Jilan Xu, Carl Thomé, Danijela Horak +2

Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentall…

cs.CV2025

Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"

Bruno Korbar, Andrew Zisserman

The goal of this paper is to be able to retrieve images using a compound query that combines object instance information from an image, with a natural text description of what that…

cs.CV2025

Understanding Co-speech Gestures in-the-wild

Sindhu B Hegde, K R Prajwal, Taein Kwon +1

Co-speech gestures play a vital role in non-verbal communication. In this paper, we introduce a new framework for co-speech gesture understanding in the wild. Specifically, we prop…

eess.AS2025

VoiceVector: Multimodal Enrolment Vectors for Speaker Separation

Akam Rahimi, Triantafyllos Afouras, Andrew Zisserman

We present a transformer-based architecture for voice separation of a target speaker from multiple other speakers and ambient noise. We achieve this by using two separate neural ne…

cs.CV2024

New keypoint-based approach for recognising British Sign Language (BSL) from sequences

Oishi Deb, KR Prajwal, Andrew Zisserman

In this paper, we present a novel keypoint-based classification model designed to recognise British Sign Language (BSL) words within continuous signing sequences. Our model's perfo…