collaborators

5 papers

cs.CV2025

Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval

Adriano Fragomeni, Dima Damen, Michael Wray

Video retrieval requires aligning visual content with corresponding natural language descriptions. In this paper, we introduce Modality Auxiliary Concepts for Video Retrieval (MAC-…

cs.CV2025

Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review

Adriano Fragomeni, Dima Damen, Michael Wray

Text-to-Video (T2V) retrieval aims to identify the most relevant item from a gallery of videos based on a user's text query. Traditional methods rely solely on aligning video and t…

cs.CV2025

HD-EPIC: A Highly-Detailed Egocentric Video Dataset

Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha +16

We present a validation dataset of newly-collected kitchen-based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe…

cs.CV2025

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

Tomáš Souček, Prajwal Gatti, Michael Wray +3

The goal of this work is to generate step-by-step visual instructions in the form of a sequence of images, given an input image that provides the scene context and the sequence of…

cs.CV2025

Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval

Kevin Flanagan, Dima Damen, Michael Wray

Video Moment Retrieval is a common task to evaluate the performance of visual-language models - it involves localising start and end times of moments in videos from query sentences…