activity
20162023
most citedLearning to Inpaint for Image Compression

37 citations · 172 across the 28 of their papers we have counts for

collaborators
Showing 2021 · cs.CVShow all

6 papers · 2 filters

cs.CV2021

Label Hallucination for Few-Shot Classification

Yiren Jian, Lorenzo Torresani

Few-shot classification requires adapting knowledge learned from a large annotated base dataset to recognize novel unseen classes, each represented by few labeled examples. In such…

cs.CV2021★ 7 cited

Ego4D: Around the World in 3,000 Hours of Egocentric Video

Kristen Grauman, Andrew Westbury, Eugene Byrne +82

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outd…

cs.CV2021

Long-Short Temporal Contrastive Learning of Video Transformers

Jue Wang, Gedas Bertasius, Du Tran +1

Video transformers have recently emerged as a competitive alternative to 3D CNNs for video understanding. However, due to their large number of parameters and reduced inductive bia…

cs.CV2021

Beyond Short Clips: End-to-End Video-Level Learning with Collaborative Memories

Xitong Yang, Haoqi Fan, Lorenzo Torresani +2

The standard way of training video models entails sampling at each iteration a single clip from a video and optimizing the clip prediction with respect to the video-level label. We…

cs.CV2021

Is Space-Time Attention All You Need for Video Understanding?

Gedas Bertasius, Heng Wang, Lorenzo Torresani

We present a convolution-free approach to video classification built exclusively on self-attention over space and time. Our method, named "TimeSformer," adapts the standard Transfo…

cs.CV2021★ 8 cited

VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs

Xudong Lin, Gedas Bertasius, Jue Wang +3

We present \textsc{Vx2Text}, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, whic…