activity
20212023
most citedMask2Former for Video Instance Segmentation

64 citations · 98 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV202314 cited

ImageBind: One Embedding Space To Bind Them All

Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu +4

We present ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. We show that all combinations of p…

cs.CV202311 cited

Cut and Learn for Unsupervised Object Detection and Instance Segmentation

Xudong Wang, Rohit Girdhar, Stella X. Yu +1

We propose Cut-and-LEaRn (CutLER), a simple approach for training unsupervised object detection and segmentation models. We leverage the property of self-supervised models to 'disc…

cs.CV20234 cited

Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities

Melissa Hall, Laura Gustafson, Aaron Adcock +2

We explore the extent to which zero-shot vision-language models exhibit gender bias for different vision tasks. Vision models traditionally required task-specific labels for repres…

cs.CV20233 cited

A Simple Recipe for Competitive Low-compute Self supervised Vision Models

Quentin Duval, Ishan Misra, Nicolas Ballas

Self-supervised methods in vision have been mostly focused on large architectures as they seem to suffer from a significant performance drop for smaller architectures. In this pape…

cs.CV2022

Omnivore: A Single Model for Many Visual Modalities

Rohit Girdhar, Mannat Singh, Nikhila Ravi +3

Prior work has studied different visual modalities in isolation and developed separate architectures for recognition of images, videos, and 3D data. Instead, in this paper, we prop…

cs.CV202164 cited

Mask2Former for Video Instance Segmentation

Bowen Cheng, Anwesa Choudhuri, Ishan Misra +3

We find Mask2Former also achieves state-of-the-art performance on video instance segmentation without modifying the architecture, the loss or even the training pipeline. In this re…