6 papers
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
Sethuraman TV, Savya Khosla, Vignesh Srinivasakumar +5
Dense video prediction tasks, such as object tracking and semantic segmentation, require video encoders that generate temporally consistent, spatially dense features for every fram…
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
Yao Xiao, Qiqian Fu, Heyi Tao +3
Image-text models excel at image-level tasks but struggle with detailed visual understanding. While these models provide strong visual-language alignment, segmentation models like…
PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
Ansel Blume, Jeonghwan Kim, Hyeonjeong Ha +7
Real-world objects are composed of distinctive, object-specific parts. Identifying these parts is key to performing fine-grained, compositional reasoning-yet, large multimodal mode…
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
Savya Khosla, Sethuraman TV, Barnett Lee +2
We introduce the Region Encoder Network (REN), a fast and effective model for generating region-based image representations using point prompts. Recent methods combine class-agnost…
Visual Program Distillation with Template-Based Augmentation
Michal Shlapentokh-Rothman, Yu-Xiong Wang, Derek Hoiem
Adapting visual programming or prompting large language models (LLMs) to generate executable code for visual tasks like visual question answering (VQA) for specialized tasks or dom…
Plenoptic PNG: Real-Time Neural Radiance Fields in 150 KB
Jae Yong Lee, Yuqun Wu, Chuhang Zou +2
The goal of this paper is to encode a 3D scene into an extremely compact representation from 2D images and to enable its transmittance, decoding and rendering in real-time across v…