3 papers
cs.CV2025
Large-scale Pre-training for Grounded Video Caption Generation
Evangelos Kazakos, Cordelia Schmid, Josef Sivic
We propose a novel approach for captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally dense bounding boxes. We introdu…
cs.CV2025
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
Tomáš SouÄek, Prajwal Gatti, Michael Wray +3
The goal of this work is to generate step-by-step visual instructions in the form of a sequence of images, given an input image that provides the scene context and the sequence of…
cs.CV2024
Grounded Video Caption Generation
Evangelos Kazakos, Cordelia Schmid, Josef Sivic
We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are gro…