activity
20152026
most citedHow Much Can CLIP Benefit Vision-and-Language Tasks?

153 citations · 295 across the 26 of their papers we have counts for

collaborators
Showing 2022Show all

11 papers · 1 filter

cs.CV2022★ 1 cited

Focus! Relevant and Sufficient Context Selection for News Image Captioning

Mingyang Zhou, Grace Luo, Anna Rohrbach +1

News Image Captioning requires describing an image by leveraging additional context from a news article. Previous works only coarsely leverage the article to extract the necessary…

cs.CV2022★ 7 cited

Shape-Guided Diffusion with Inside-Outside Attention

Dong Huk Park, Grace Luo, Clayton Toste +5

We introduce precise object silhouette as a new form of user control in text-to-image diffusion models, which we dub Shape-Guided Diffusion. Our training-free method uses an Inside…

cs.CV2022★ 1 cited

G^3: Geolocation via Guidebook Grounding

Grace Luo, Giscard Biamby, Trevor Darrell +2

We demonstrate how language can improve geolocation: the task of predicting the location where an image was taken. Here we study explicit knowledge from human-written guidebooks th…

cs.CV2022

TL;DW? Summarizing Instructional Videos with Task Relevance & Cross-Modal Saliency

Medhini Narasimhan, Arsha Nagrani, Chen Sun +4

YouTube users looking for instructions for a specific task may spend a long time browsing content trying to find the right video that matches their needs. Creating a visual summary…

cs.CV2022★ 1 cited

Structured Video Tokens @ Ego4D PNR Temporal Localization Challenge 2022

Elad Ben-Avraham, Roei Herzig, Karttikeya Mangalam +5

This technical report describes the SViT approach for the Ego4D Point of No Return (PNR) Temporal Localization Challenge. We propose a learning framework StructureViT (SViT for sho…

cs.CV2022★ 8 cited

Bringing Image Scene Structure to Video via Frame-Clip Consistency of Object Tokens

Elad Ben-Avraham, Roei Herzig, Karttikeya Mangalam +5

Recent action recognition models have achieved impressive results by integrating objects, their locations and interactions. However, obtaining dense structured annotations for each…