activity
20182026
most citedGRIT: General Robust Image Task Benchmark

11 citations · 18 across the 2 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

Tanmay Gupta, Piper Wolters, Zixian Ma +13

Web agents--autonomous systems that navigate and execute tasks on the web on behalf of users--have the potential to transform how people interact with the digital world. However, t…

cs.CV20227 cited

Visual Programming: Compositional visual reasoning without training

Tanmay Gupta, Aniruddha Kembhavi

We present VISPROG, a neuro-symbolic approach to solving complex and compositional visual tasks given natural language instructions. VISPROG avoids the need for any task-specific t…

cs.CV202211 cited

GRIT: General Robust Image Task Benchmark

Tanmay Gupta, Ryan Marten, Aniruddha Kembhavi +1

Computer vision models excel at making predictions when the test distribution closely resembles the training distribution. Such models have yet to match the ability of biological v…

cs.CV2021

Visual Semantic Role Labeling for Video Understanding

Arka Sadhu, Tanmay Gupta, Mark Yatskar +2

We propose a new framework for understanding and representing related salient events in a video using visual semantic role labeling. We represent videos as a set of related events,…

cs.CV2020

Contrastive Learning for Weakly Supervised Phrase Grounding

Tanmay Gupta, Arash Vahdat, Gal Chechik +3

Phrase grounding, the problem of associating image regions to caption words, is a crucial component of vision-language tasks. We show that phrase grounding can be learned by optimi…

cs.CV2019

ViCo: Word Embeddings from Visual Co-occurrences

Tanmay Gupta, Alexander Schwing, Derek Hoiem

We propose to learn word embeddings from visual co-occurrences. Two words co-occur visually if both words apply to the same image or image region. Specifically, we extract four typ…