activity
20162025
most citedClick Carving: Segmenting Objects in Video with Point Clicks

14 citations · 14 across the 2 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi +26

Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The researc…

cs.CV2024

Collecting Consistently High Quality Object Tracks with Minimal Human Involvement by Using Self-Supervised Learning to Detect Tracker Errors

Samreen Anjum, Suyog Jain, Danna Gurari

We propose a hybrid framework for consistently producing high-quality object tracks by combining an automated object tracker with little human input. The key idea is to tailor a mo…

cs.CV2023

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98

We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric…

cs.CV2019

Predicting How to Distribute Work Between Algorithms and Humans to Segment an Image Batch

Danna Gurari, Yinan Zhao, Suyog Dutt Jain +2

Foreground object segmentation is a critical step for many image analysis tasks. While automated methods can produce high-quality results, their failures disappoint users in need o…

cs.CV2018

Pixel Objectness: Learning to Segment Generic Objects Automatically in Images and Videos

Bo Xiong, Suyog Dutt Jain, Kristen Grauman

We propose an end-to-end learning framework for segmenting generic objects in both images and videos. Given a novel image or video, our approach produces a pixel-level mask for all…

cs.CV201614 cited

Click Carving: Segmenting Objects in Video with Point Clicks

Suyog Dutt Jain, Kristen Grauman

We present a novel form of interactive video object segmentation where a few clicks by the user helps the system produce a full spatio-temporal segmentation of the object of intere…