3 papers
cs.CV2025
A Video Is Not Worth a Thousand Words
Sam Pollard, Michael Wray
As we become increasingly dependent on vision language models (VLMs) to answer questions about the world around us, there is a significant amount of research devoted to increasing…
cs.CV2025
Video, How Do Your Tokens Merge?
Sam Pollard, Michael Wray
Video transformer models require huge amounts of compute resources due to the spatio-temporal scaling of the input. Tackling this, recent methods have proposed to drop or merge tok…
cs.CV2025
HD-EPIC: A Highly-Detailed Egocentric Video Dataset
Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha +16
We present a validation dataset of newly-collected kitchen-based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe…