activity
20182025
most citedVideo Understanding as Machine Translation

26 citations · 28 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2025

Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"

Bruno Korbar, Andrew Zisserman

The goal of this paper is to be able to retrieve images using a compound query that combines object instance information from an image, with a natural text description of what that…

cs.CV20241 cited

Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling

Bruno Korbar, Jaesung Huh, Andrew Zisserman

The goal of this paper is automatic character-aware subtitle generation. Given a video and a minimal amount of metadata, we propose an audio-visual method that generates a full tra…

cs.CV2023

Text-Conditioned Resampler For Long Form Video Understanding

Bruno Korbar, Yongqin Xian, Alessio Tonioni +2

In this paper we present a text-conditioned video resampler (TCR) module that uses a pre-trained and frozen visual encoder and large language model (LLM) to process long video sequ…

cs.CV20222 cited

End-to-end Tracking with a Multi-query Transformer

Bruno Korbar, Andrew Zisserman

Multiple-object tracking (MOT) is a challenging task that requires simultaneous reasoning about location, appearance, and identity of the objects in the scene over time. Our aim in…

cs.CV202026 cited

Video Understanding as Machine Translation

Bruno Korbar, Fabio Petroni, Rohit Girdhar +1

With the advent of large-scale multimodal video datasets, especially sequences with audio or transcribed speech, there has been a growing interest in self-supervised learning of vi…

cs.CV2019

SCSampler: Sampling Salient Clips from Video for Efficient Action Recognition

Bruno Korbar, Du Tran, Lorenzo Torresani

While many action recognition datasets consist of collections of brief, trimmed videos each containing a relevant action, videos in the real-world (e.g., on YouTube) exhibit very d…