26 citations · 28 across the 3 of their papers we have counts for
7 papers · 1 filter
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
Bruno Korbar, Andrew Zisserman
The goal of this paper is to be able to retrieve images using a compound query that combines object instance information from an image, with a natural text description of what that…
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
Bruno Korbar, Jaesung Huh, Andrew Zisserman
The goal of this paper is automatic character-aware subtitle generation. Given a video and a minimal amount of metadata, we propose an audio-visual method that generates a full tra…
Text-Conditioned Resampler For Long Form Video Understanding
Bruno Korbar, Yongqin Xian, Alessio Tonioni +2
In this paper we present a text-conditioned video resampler (TCR) module that uses a pre-trained and frozen visual encoder and large language model (LLM) to process long video sequ…
End-to-end Tracking with a Multi-query Transformer
Bruno Korbar, Andrew Zisserman
Multiple-object tracking (MOT) is a challenging task that requires simultaneous reasoning about location, appearance, and identity of the objects in the scene over time. Our aim in…
Video Understanding as Machine Translation
Bruno Korbar, Fabio Petroni, Rohit Girdhar +1
With the advent of large-scale multimodal video datasets, especially sequences with audio or transcribed speech, there has been a growing interest in self-supervised learning of vi…
SCSampler: Sampling Salient Clips from Video for Efficient Action Recognition
Bruno Korbar, Du Tran, Lorenzo Torresani
While many action recognition datasets consist of collections of brief, trimmed videos each containing a relevant action, videos in the real-world (e.g., on YouTube) exhibit very d…