Publications (29)
Open-world Text-specified Object Counting
Niki Amini-Naieni, Kiana Amini-Naieni, Tengda Han +1
Our objective is open-world object counting in images, where the target object class is specified by a text description. To this end, we propose CounTX, a class-agnostic, single-st…
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc +24
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Fl…
WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
Max Bain, Jaesung Huh, Tengda Han +1
Large-scale, weakly-supervised speech recognition models, such as Whisper, have demonstrated impressive results on speech recognition across domains and languages. However, their a…
Human Pose Forecasting via Deep Markov Models
Sam Toyer, Anoop Cherian, Tengda Han +1
Human pose forecasting is an important problem in computer vision with applications to human-robot interaction, visual surveillance, and autonomous driving. Usually, forecasting al…
Human Action Forecasting by Learning Task Grammars
Tengda Han, Jue Wang, Anoop Cherian +1
For effective human-robot interaction, it is important that a robotic assistant can forecast the next action a human will consider in a given task. Unfortunately, real-world tasks…
Self-supervised Co-training for Video Representation Learning
Tengda Han, Weidi Xie, Andrew Zisserman
The objective of this paper is visual-only self-supervised video representation learning. We make the following contributions: (i) we investigate the benefit of adding semantic-cla…
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
Toby Perrett, Tengda Han, Dima Damen +1
Long videos contain many repeating actions, events and shots. These repetitions are frequently given identical captions, which makes it difficult to retrieve the exact desired clip…
Learning from Streaming Video with Orthogonal Gradients
Tengda Han, Dilara Gokay, Joseph Heyward +6
We address the challenge of representation learning from a continuous stream of video as input, in a self-supervised manner. This differs from the standard approaches to video lear…
Temporal Alignment Networks for Long-term Video
Tengda Han, Weidi Xie, Andrew Zisserman
The objective of this paper is a temporal alignment network that ingests long term video sequences, and associated text sentences, in order to: (1) determine if a sentence is align…
Turbo Training with Token Dropout
Tengda Han, Weidi Xie, Andrew Zisserman
The objective of this paper is an efficient training method for video tasks. We make three contributions: (1) We propose Turbo training, a simple and versatile training paradigm fo…
Seeing without Pixels: Perception from Camera Trajectories
Zihui Xue, Kristen Grauman, Dima Damen +2
Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigat…
Character-Centric Understanding of Animated Movies
Zhongrui Gui, Junyu Xie, Tengda Han +2
Animated movies are captivating for their unique character designs and imaginative storytelling, yet they pose significant challenges for existing recognition systems. Unlike the c…
AutoAD II: The Sequel -- Who, When, and What in Movie Audio Description
Tengda Han, Max Bain, Arsha Nagrani +3
Audio Description (AD) is the task of generating descriptions of visual content, at suitable time intervals, for the benefit of visually impaired audiences. For movies, this presen…
Learning to Count without Annotations
Lukas Knobel, Tengda Han, Yuki M. Asano
While recent supervised methods for reference-based object counting continue to improve the performance on benchmark datasets, they have to rely on small datasets due to the cost a…
Memory-augmented Dense Predictive Coding for Video Representation Learning
Tengda Han, Weidi Xie, Andrew Zisserman
The objective of this paper is self-supervised learning from video, in particular for representations for action recognition. We make the following contributions: (i) We propose a…
CountGD: Multi-Modal Open-World Counting
Niki Amini-Naieni, Tengda Han, Andrew Zisserman
The goal of this paper is to improve the generality and accuracy of open-vocabulary object counting in images. To improve the generality, we repurpose an open-vocabulary detection…
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
Junyu Xie, Tengda Han, Weidi Xie +1
Moving Object Segmentation (MOS) aims to discover, segment, and track objects that move independently of the camera. Current MOS methods, however, exhibit two fundamental limitatio…
Dynamic Reflections: Probing Video Representations with Text Alignment
Tyler Zhu, Tengda Han, Leonidas Guibas +2
The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encod…
More than a Moment: Towards Coherent Sequences of Audio Descriptions
Eshika Khandelwal, Junyu Xie, Tengda Han +5
Audio Descriptions (ADs) convey essential on-screen information, allowing visually impaired audiences to follow videos. To be effective, ADs must form a coherent sequence that help…
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
Junyu Xie, Tengda Han, Max Bain +5
Our objective is the automatic generation of Audio Descriptions (ADs) for edited video material, such as movies and TV series. To achieve this, we propose a two-stage framework tha…
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
Junyu Xie, Tengda Han, Max Bain +4
Our objective is to generate Audio Descriptions (ADs) for both movies and TV series in a training-free manner. We use the power of off-the-shelf Visual-Language Models (VLMs) and L…
Unique Lives, Shared World: Learning from Single-Life Videos
Tengda Han, Sayna Ebrahimi, Dilara Gokay +8
We introduce the "single-life" learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual. We leverage the multiple viewp…
AutoAD: Movie Description in Context
Tengda Han, Max Bain, Arsha Nagrani +3
The objective of this paper is an automatic Audio Description (AD) model that ingests movies and outputs AD in text form. Generating high-quality movie AD is challenging due to the…
AutoAD III: The Prequel -- Back to the Pixels
Tengda Han, Max Bain, Arsha Nagrani +3
Generating Audio Description (AD) for movies is a challenging task that requires fine-grained visual understanding and an awareness of the characters and their names. Currently, vi…
Prompt Generation Networks for Input-Space Adaptation of Frozen Vision Transformers
Jochem Loedeman, Maarten C. Stol, Tengda Han +1
With the introduction of the transformer architecture in computer vision, increasing model scale has been demonstrated as a clear path to achieving performance and robustness gains…
Video Representation Learning by Dense Predictive Coding
Tengda Han, Weidi Xie, Andrew Zisserman
The objective of this paper is self-supervised learning of spatio-temporal embeddings from video, suitable for human action recognition. We make three contributions: First, we intr…
Prompting Visual-Language Models for Efficient Video Understanding
Chen Ju, Tengda Han, Kunhao Zheng +2
Image-based visual-language (I-VL) pre-training has shown great success for learning joint visual-textual representations from large-scale web data, revealing remarkable ability fo…
Multi-Sentence Grounding for Long-term Instructional Video
Zeqian Li, Qirui Chen, Tengda Han +3
In this paper, we aim to establish an automatic, scalable pipeline for denoising the large-scale instructional dataset and construct a high-quality video-text dataset with multiple…
Stale Diffusion: Hyper-realistic 5D Movie Generation Using Old-school Methods
Joao F. Henriques, Dylan Campbell, Tengda Han
Two years ago, Stable Diffusion achieved super-human performance at generating images with super-human numbers of fingers. Following the steady decline of its technical novelty, we…