activity
20172025
most citedThinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers

133 citations · 431 across the 32 of their papers we have counts for

collaborators
Showing cs.CVShow all

39 papers · 1 filter

cs.CV2025

Large-scale Pre-training for Grounded Video Caption Generation

Evangelos Kazakos, Cordelia Schmid, Josef Sivic

We propose a novel approach for captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally dense bounding boxes. We introdu…

cs.CV2024

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

Tomáš Souček, Prajwal Gatti, Michael Wray +3

The goal of this work is to generate step-by-step visual instructions in the form of a sequence of images, given an input image that provides the scene context and the sequence of…

cs.CV2024★ 1 cited

Grounded Video Caption Generation

Evangelos Kazakos, Cordelia Schmid, Josef Sivic

We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are gro…

cs.CV2023

GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos

Tomáš Souček, Dima Damen, Michael Wray +2

We address the task of generating temporally consistent and physically plausible images of actions and object state transformations. Given an input image and a text prompt describi…

cs.CV2023★ 2 cited

VidChapters-7M: Video Chapters at Scale

Antoine Yang, Arsha Nagrani, Ivan Laptev +2

Segmenting long videos into chapters enables users to quickly navigate to the information of their interest. This important topic has been understudied due to the lack of publicly…

cs.CV2023

Meta-Personalizing Vision-Language Models to Find Named Instances in Video

Chun-Hsiao Yeh, Bryan Russell, Josef Sivic +2

Large-scale vision-language models (VLM) have shown impressive results for language-guided search applications. While these models allow category-level queries, they currently stru…