4 papers
Beyond Caption-Based Queries for Video Moment Retrieval
David Pujol-Perich, Albert Clapés, Dima Damen +2
In this work, we investigate the degradation of existing VMR methods, particularly of DETR architectures, when trained on caption-based queries but evaluated on search queries. For…
A Video Is Not Worth a Thousand Words
Sam Pollard, Michael Wray
As we become increasingly dependent on vision language models (VLMs) to answer questions about the world around us, there is a significant amount of research devoted to increasing…
Evaluating Compositional Generalisation in VLMs and Diffusion Models
Beth Pearson, Bilal Boulbarss, Michael Wray +1
A fundamental aspect of the semantics of natural language is that novel meanings can be formed from the composition of previously known parts. Vision-language models (VLMs) have ma…
Video, How Do Your Tokens Merge?
Sam Pollard, Michael Wray
Video transformer models require huge amounts of compute resources due to the spatio-temporal scaling of the input. Tackling this, recent methods have proposed to drop or merge tok…