3 citations · 4 across the 14 of their papers we have counts for
15 papers · 1 filter
ViTAL-X: Video-Text Alignment with Cross-Modal Temporal Edits
Sethuraman T, Savya Khosla, Onkar Kishor Susladkar +6
Video-text models adapted from image-text architectures (e.g., CLIP) frequently exhibit temporal blindness, the inability to perceive fundamental cues like order, direction, and mo…
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Yao Xiao, Reuben Tan, Zhen Zhu +3
Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infe…
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
Michal Shlapentokh-Rothman, Prachi Garg, Yu-Xiong Wang +1
Keyframe selection is a direct way to provide verifiable visual evidence for long-video question answering (QA). Queries differ in what they require, and finding the right frames d…
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
Savya Khosla, Sethuraman T, Aryan Chadha +2
Despite recent progress, vision-language encoders struggle with two core limitations: (1) weak alignment between language and dense vision features, which hurts tasks like open-voc…
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
Sethuraman T, Savya Khosla, Aditi Tiwari +11
This work investigates a fundamental question: Do Video-Language Models (VidLMs) robustly account for video content, temporal sequence, and motion? Our investigation shows that, su…
SceneDiff: A Benchmark and Method for Multiview Object Change Detection
Yuqun Wu, Chih-hao Lin, Henry Che +4
We investigate the problem of identifying objects that have been added, removed, or moved between a pair of captures (images or videos) of the same scene at different times. Accura…