activity
20192025
most citedMeasuring Robustness to Natural Distribution Shifts in Image Classification

170 citations · 175 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2025

AllTracker: Efficient Dense Point Tracking at High Resolution

Adam W. Harley, Yang You, Xinglong Sun +11

We introduce AllTracker: a model that estimates long-range point tracks by way of estimating the flow field between a query frame and every other frame of a video. Unlike existing…

cs.CV2025

Understanding Complexity in VideoQA via Visual Program Generation

Cristobal Eyzaguirre, Igor Vasiljevic, Achal Dave +5

We propose a data-driven approach to analyzing query complexity in Video Question Answering (VideoQA). Previous efforts in benchmark design have relied on human expertise to design…

cs.CV2025

Should VLMs be Pre-trained with Image Data?

Sedrick Keh, Jean Mercat, Samir Yitzhak Gadre +8

Pre-trained LLMs that are further trained with image data perform well on vision-language tasks. While adding images during a second training phase effectively unlocks this capabil…

cs.CV2024

Espresso: High Compression For Rich Extraction From Videos for Your Vision-Language Model

Keunwoo Peter Yu, Achal Dave, Rares Ambrus +1

Recent advances in vision-language models (VLMs) have shown great promise in connecting images and text, but extending these models to long videos remains challenging due to the ra…

cs.CV2024

GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusion

Vitor Guizilini, Pavel Tokmakov, Achal Dave +1

3D reconstruction from a single image is a long-standing problem in computer vision. Learning-based methods address its inherent scale ambiguity by leveraging increasingly large la…

cs.CV20221 cited

Differentiable Raycasting for Self-supervised Occupancy Forecasting

Tarasha Khurana, Peiyun Hu, Achal Dave +3

Motion planning for safe autonomous driving requires learning how the environment around an ego-vehicle evolves with time. Ego-centric perception of driveable regions in a scene no…