7 papers · 1 filter
AllTracker: Efficient Dense Point Tracking at High Resolution
Adam W. Harley, Yang You, Xinglong Sun +11
We introduce AllTracker: a model that estimates long-range point tracks by way of estimating the flow field between a query frame and every other frame of a video. Unlike existing…
Understanding Complexity in VideoQA via Visual Program Generation
Cristobal Eyzaguirre, Igor Vasiljevic, Achal Dave +5
We propose a data-driven approach to analyzing query complexity in Video Question Answering (VideoQA). Previous efforts in benchmark design have relied on human expertise to design…
Espresso: High Compression For Rich Extraction From Videos for Your Vision-Language Model
Keunwoo Peter Yu, Achal Dave, Rares Ambrus +1
Recent advances in vision-language models (VLMs) have shown great promise in connecting images and text, but extending these models to long videos remains challenging due to the ra…
Should VLMs be Pre-trained with Image Data?
Sedrick Keh, Jean Mercat, Samir Yitzhak Gadre +8
Pre-trained LLMs that are further trained with image data perform well on vision-language tasks. While adding images during a second training phase effectively unlocks this capabil…
GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusion
Vitor Guizilini, Pavel Tokmakov, Achal Dave +1
3D reconstruction from a single image is a long-standing problem in computer vision. Learning-based methods address its inherent scale ambiguity by leveraging increasingly large la…
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis
Basile Van Hoorick, Rundi Wu, Ege Ozguroglu +6
Accurate reconstruction of complex dynamic scenes from just a single viewpoint continues to be a challenging task in computer vision. Current dynamic novel view synthesis methods t…