6 papers
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
Savya Khosla, Sethuraman T, Aryan Chadha +2
Despite recent progress, vision-language encoders struggle with two core limitations: (1) weak alignment between language and dense vision features, which hurts tasks like open-voc…
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
Sethuraman T, Savya Khosla, Aditi Tiwari +11
This work investigates a fundamental question: Do Video-Language Models (VidLMs) robustly account for video content, temporal sequence, and motion? Our investigation shows that, su…
RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
Savya Khosla, Sethuraman T, Alexander Schwing +1
We present RELOCATE, a simple training-free baseline designed to perform the challenging task of visual query localization in long videos. To eliminate the need for task-specific t…
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
Savya Khosla, Sethuraman TV, Barnett Lee +2
We introduce the Region Encoder Network (REN), a fast and effective model for generating region-based image representations using point prompts. Recent methods combine class-agnost…
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
Sethuraman TV, Savya Khosla, Vignesh Srinivasakumar +5
Dense video prediction tasks, such as object tracking and semantic segmentation, require video encoders that generate temporally consistent, spatially dense features for every fram…
MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities
Savya Khosla, Aditi Tiwari, Kushal Kafle +4
While originally designed for unidirectional generative modeling, decoder-only large language models (LLMs) are increasingly being adapted for bidirectional modeling. However, unid…