activity
20172023
most citedLearning to Inpaint for Image Compression

37 citations · 140 across the 17 of their papers we have counts for

collaborators
Showing cs.CVShow all

27 papers · 1 filter

cs.CV2023

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98

We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric…

cs.CV2023

Multiscale Video Pretraining for Long-Term Activity Forecasting

Reuben Tan, Matthias De Lange, Michael Iuzzolino +4

Long-term activity forecasting is an especially challenging research problem because it requires understanding the temporal relationships between observed actions, as well as the v…

cs.CV20231 cited

Learning to Ground Instructional Articles in Videos through Narrations

Effrosyni Mavroudi, Triantafyllos Afouras, Lorenzo Torresani

In this paper we present an approach for localizing steps of procedural activities in narrated how-to videos. To deal with the scarcity of labeled data at scale, we source the step…

cs.CV20231 cited

MINOTAUR: Multi-task Video Grounding From Multimodal Queries

Raghav Goyal, Effrosyni Mavroudi, Xitong Yang +5

Video understanding tasks take many forms, from action detection to visual query localization and spatio-temporal grounding of sentences. These tasks differ in the type of inputs (…

cs.CV2023

Open-world Instance Segmentation: Top-down Learning with Bottom-up Supervision

Tarun Kalluri, Weiyao Wang, Heng Wang +3

Many top-down architectures for instance segmentation achieve significant success when trained and tested on pre-defined closed-world taxonomy. However, when deployed in the open w…

cs.CV2022

Deformable Video Transformer

Jue Wang, Lorenzo Torresani

Video transformers have recently emerged as an effective alternative to convolutional networks for action classification. However, most prior video transformers adopt either global…