activity
20232026
most citedUnified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

3 citations · 4 across the 14 of their papers we have counts for

collaborators
Showing cs.CVShow all

15 papers · 1 filter

cs.CV2026

ViTAL-X: Video-Text Alignment with Cross-Modal Temporal Edits

Sethuraman T, Savya Khosla, Onkar Kishor Susladkar +6

Video-text models adapted from image-text architectures (e.g., CLIP) frequently exhibit temporal blindness, the inability to perceive fundamental cues like order, direction, and mo…

cs.CV2026

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Yao Xiao, Reuben Tan, Zhen Zhu +3

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infe…

cs.CV2026

Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval

Michal Shlapentokh-Rothman, Prachi Garg, Yu-Xiong Wang +1

Keyframe selection is a direct way to provide verifiable visual evidence for long-video question answering (QA). Queries differ in what they require, and finding the right frames d…

cs.CV2026

T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability

Savya Khosla, Sethuraman T, Aryan Chadha +2

Despite recent progress, vision-language encoders struggle with two core limitations: (1) weak alignment between language and dense vision features, which hurts tasks like open-voc…

cs.CV2026

Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models

Sethuraman T, Savya Khosla, Aditi Tiwari +11

This work investigates a fundamental question: Do Video-Language Models (VidLMs) robustly account for video content, temporal sequence, and motion? Our investigation shows that, su…

cs.CV2025

SceneDiff: A Benchmark and Method for Multiview Object Change Detection

Yuqun Wu, Chih-hao Lin, Henry Che +4

We investigate the problem of identifying objects that have been added, removed, or moved between a pair of captures (images or videos) of the same scene at different times. Accura…