13 citations · 16 across the 6 of their papers we have counts for
6 papers
Semantic Compositions Enhance Vision-Language Contrastive Learning
Maxwell Aladago, Lorenzo Torresani, Soroush Vosoughi
In the field of vision-language contrastive learning, models such as CLIP capitalize on matched image-caption pairs as positive examples and leverage within-batch non-matching pair…
Multiscale Video Pretraining for Long-Term Activity Forecasting
Reuben Tan, Matthias De Lange, Michael Iuzzolino +4
Long-term activity forecasting is an especially challenging research problem because it requires understanding the temporal relationships between observed actions, as well as the v…
Learning to Ground Instructional Articles in Videos through Narrations
Effrosyni Mavroudi, Triantafyllos Afouras, Lorenzo Torresani
In this paper we present an approach for localizing steps of procedural activities in narrated how-to videos. To deal with the scarcity of labeled data at scale, we source the step…
MINOTAUR: Multi-task Video Grounding From Multimodal Queries
Raghav Goyal, Effrosyni Mavroudi, Xitong Yang +5
Video understanding tasks take many forms, from action detection to visual query localization and spatio-temporal grounding of sentences. These tasks differ in the type of inputs (…
Egocentric Video Task Translation @ Ego4D Challenge 2022
Zihui Xue, Yale Song, Kristen Grauman +1
This technical report describes the EgoTask Translation approach that explores relations among a set of egocentric video tasks in the Ego4D challenge. To improve the primary task o…
DeepEdge: A Multi-Scale Bifurcated Deep Network for Top-Down Contour Detection
Gedas Bertasius, Jianbo Shi, Lorenzo Torresani
Contour detection has been a fundamental component in many image segmentation and object detection systems. Most previous work utilizes low-level features such as texture or salien…