630 citations · 2.3k across the 46 of their papers we have counts for
12 papers · 1 filter
Multitask Vision-Language Prompt Tuning
Sheng Shen, Shijia Yang, Tianjun Zhang +4
Prompt Tuning, conditioning on task-specific learned prompt vectors, has emerged as a data-efficient and parameter-efficient method for adapting large pretrained vision-language mo…
G^3: Geolocation via Guidebook Grounding
Grace Luo, Giscard Biamby, Trevor Darrell +2
We demonstrate how language can improve geolocation: the task of predicting the location where an image was taken. Here we study explicit knowledge from human-written guidebooks th…
Real-World Robot Learning with Masked Visual Pre-training
Ilija Radosavovic, Tete Xiao, Stephen James +3
In this work, we explore self-supervised visual pre-training on images from diverse, in-the-wild videos for real-world robotic tasks. Like prior work, our visual representations ar…
Studying Bias in GANs through the Lens of Race
Vongani H. Maluleke, Neerja Thakkar, Tim Brooks +5
In this work, we study how the performance and evaluation of generative image models are impacted by the racial composition of their training datasets. By examining and controlling…
Prior Knowledge-Guided Attention in Self-Supervised Vision Transformers
Kevin Miao, Akash Gokul, Raghav Singh +5
Recent trends in self-supervised representation learning have focused on removing inductive biases from training pipelines. However, inductive biases can be useful in settings when…
Voxel-informed Language Grounding
Rodolfo Corona, Shizhan Zhu, Dan Klein +1
Natural language applied to natural 2D images describes a fundamentally 3D world. We present the Voxel-informed Language Grounder (VLG), a language grounding model that leverages 3…