4 papers
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
Savya Khosla, Sethuraman T, Aryan Chadha +2
Despite recent progress, vision-language encoders struggle with two core limitations: (1) weak alignment between language and dense vision features, which hurts tasks like open-voc…
RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
Savya Khosla, Sethuraman T, Alexander Schwing +1
We present RELOCATE, a simple training-free baseline designed to perform the challenging task of visual query localization in long videos. To eliminate the need for task-specific t…
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
Savya Khosla, Sethuraman TV, Barnett Lee +2
We introduce the Region Encoder Network (REN), a fast and effective model for generating region-based image representations using point prompts. Recent methods combine class-agnost…
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
Sethuraman TV, Savya Khosla, Vignesh Srinivasakumar +5
Dense video prediction tasks, such as object tracking and semantic segmentation, require video encoders that generate temporally consistent, spatially dense features for every fram…