16 citations · 16 across the 4 of their papers we have counts for
7 papers · 1 filter
Pano3D: Unified 3D Reconstruction and Panoptic Segmentation
Victor Barberteguy, Ahmet Iscen, Mathilde Caron +3
Recent advances in 3D feedforward reconstruction neural networks have achieved remarkable success in dense reconstruction from images without any camera parameters. Yet, equipping…
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
Lucas Ventura, Antoine Yang, Cordelia Schmid +1
We address the task of video chaptering, i.e., partitioning a long video timeline into semantic units and generating corresponding chapter titles. While relatively underexplored, a…
Learning text-to-video retrieval from image captioning
Lucas Ventura, Cordelia Schmid, Gül Varol
We describe a protocol to study text-to-video retrieval training with unlabeled videos, where we assume (i) no access to labels for any videos, i.e., no access to the set of ground…
CoVR-2: Automatic Data Construction for Composed Video Retrieval
Lucas Ventura, Antoine Yang, Cordelia Schmid +1
Composed Image Retrieval (CoIR) has recently gained popularity as a task that considers both text and image queries together, to search for relevant images in a database. Most CoIR…
Scaling up sign spotting through sign language dictionaries
Gül Varol, Liliane Momeni, Samuel Albanie +2
The focus of this work is - given a video of an isolated sign, our task is to identify and it has been signed in a cont…
Learning joint reconstruction of hands and manipulated objects
Yana Hasson, Gül Varol, Dimitrios Tzionas +4
Estimating hand-object manipulations is essential for interpreting and imitating human actions. Previous work has made significant progress towards reconstruction of hand poses and…