activity
20152022
most citedHow Much Can CLIP Benefit Vision-and-Language Tasks?

153 citations · 187 across the 9 of their papers we have counts for

collaborators

22 papers

cs.CV20221 cited

Focus! Relevant and Sufficient Context Selection for News Image Captioning

Mingyang Zhou, Grace Luo, Anna Rohrbach +1

News Image Captioning requires describing an image by leveraging additional context from a news article. Previous works only coarsely leverage the article to extract the necessary…

cs.CV20221 cited

G^3: Geolocation via Guidebook Grounding

Grace Luo, Giscard Biamby, Trevor Darrell +2

We demonstrate how language can improve geolocation: the task of predicting the location where an image was taken. Here we study explicit knowledge from human-written guidebooks th…

cs.CV20223 cited

ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension

Sanjay Subramanian, William Merrill, Trevor Darrell +3

Training a referring expression comprehension (ReC) model for a new visual domain requires collecting referring expressions, and potentially corresponding bounding boxes, for image…

cs.CV2022

On Guiding Visual Attention with Language Specification

Suzanne Petryk, Lisa Dunlap, Keyan Nasseri +3

While real world challenges typically define visual categories with language words or phrases, most visual classification methods define categories with numerical indices. However,…

cs.CV2021153 cited

How Much Can CLIP Benefit Vision-and-Language Tasks?

Sheng Shen, Liunian Harold Li, Hao Tan +5

Most existing Vision-and-Language (V&L) models rely on pre-trained visual encoders, using a relatively small set of manually-annotated data (as compared to web-crawled data), to pe…

cs.CV2021

NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal Media

Grace Luo, Trevor Darrell, Anna Rohrbach

Online misinformation is a prevalent societal issue, with adversaries relying on tools ranging from cheap fakes to sophisticated deep fakes. We are motivated by the threat scenario…