348 citations · 372 across the 5 of their papers we have counts for
15 papers · 1 filter
Open-World Instance Segmentation: Exploiting Pseudo Ground Truth From Learned Pairwise Affinity
Weiyao Wang, Matt Feiszli, Heng Wang +2
Open-world instance segmentation is the task of grouping pixels into object instances without any pre-determined taxonomy. This is challenging, as state-of-the-art methods rely on…
Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation
Weiyao Wang, Matt Feiszli, Heng Wang +1
Current state-of-the-art object detection and segmentation methods work well under the closed-world assumption. This closed-world setting assumes that the list of object categories…
Self-Supervised Learning by Cross-Modal Audio-Video Clustering
Humam Alwassel, Dhruv Mahajan, Bruno Korbar +3
Visual and audio modalities are highly correlated, yet they contain different information. Their strong correlation makes it possible to predict the semantics of one from the other…
UniDual: A Unified Model for Image and Video Understanding
Yufei Wang, Du Tran, Lorenzo Torresani
Although a video is effectively a sequence of images, visual perception systems typically model images and videos separately, thus failing to exploit the correlation and the synerg…
FASTER Recurrent Networks for Efficient Video Classification
Linchao Zhu, Laura Sevilla-Lara, Du Tran +3
Typical video classification methods often divide a video into short clips, do inference on each clip independently, then aggregate the clip-level predictions to generate the video…
Learning Temporal Pose Estimation from Sparsely-Labeled Videos
Gedas Bertasius, Christoph Feichtenhofer, Du Tran +2
Modern approaches for multi-person pose estimation in video require large amounts of dense annotations. However, labeling every frame in a video is costly and labor intensive. To r…