collaborators

5 papers

cs.CV2026

Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition

Prajwal Gatti, Simon Jenni, Fabian Caba Heilbron +1

We address the problem of training on long-tailed data for video action recognition. We propose to augment the training set using a text-to-video generative model, conditioned on d…

cs.CV2026

The Indra Representation Hypothesis for Multimodal Alignment

Jianglin Lu, Hailing Wang, Kuo Yang +3

Recent studies have uncovered an interesting phenomenon: unimodal foundation models tend to learn convergent representations, regardless of differences in architecture, training ob…

cs.CV2026

Seeing Through Words: Controlling Visual Retrieval Quality with Language Models

Jianglin Lu, Simon Jenni, Kushal Kafle +3

Text-to-image retrieval is a fundamental task in vision-language learning, yet in real-world scenarios it is often challenged by short and underspecified user queries. Such queries…

cs.CV2026

Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models

Sethuraman T, Savya Khosla, Aditi Tiwari +11

This work investigates a fundamental question: Do Video-Language Models (VidLMs) robustly account for video content, temporal sequence, and motion? Our investigation shows that, su…

cs.CV2025

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory

Sethuraman TV, Savya Khosla, Vignesh Srinivasakumar +5

Dense video prediction tasks, such as object tracking and semantic segmentation, require video encoders that generate temporally consistent, spatially dense features for every fram…