activity
20212023
most citedMask2Former for Video Instance Segmentation

64 citations · 102 across the 8 of their papers we have counts for

collaborators

8 papers

cs.RO20234 cited

On Bringing Robots Home

Nur Muhammad Mahi Shafiullah, Anant Rai, Haritheja Etukuru +4

Throughout history, we have successfully integrated various machines into our homes. Dishwashers, laundry machines, stand mixers, and robot vacuums are a few recent examples. Howev…

cs.CV202314 cited

ImageBind: One Embedding Space To Bind Them All

Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu +4

We present ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. We show that all combinations of p…

cs.LG20232 cited

RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data

Sangwoo Mo, Jong-Chyi Su, Chih-Yao Ma +4

Semi-supervised learning aims to train a model using limited labels. State-of-the-art semi-supervised methods for image classification such as PAWS rely on self-supervised represen…

cs.CV202311 cited

Cut and Learn for Unsupervised Object Detection and Instance Segmentation

Xudong Wang, Rohit Girdhar, Stella X. Yu +1

We propose Cut-and-LEaRn (CutLER), a simple approach for training unsupervised object detection and segmentation models. We leverage the property of self-supervised models to 'disc…

cs.CV20234 cited

Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities

Melissa Hall, Laura Gustafson, Aaron Adcock +2

We explore the extent to which zero-shot vision-language models exhibit gender bias for different vision tasks. Vision models traditionally required task-specific labels for repres…

cs.CV20233 cited

A Simple Recipe for Competitive Low-compute Self supervised Vision Models

Quentin Duval, Ishan Misra, Nicolas Ballas

Self-supervised methods in vision have been mostly focused on large architectures as they seem to suffer from a significant performance drop for smaller architectures. In this pape…