Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
Sirnam Swetha, Rohit Gupta, Parth Parag Kulkarni +5
Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly…
cs.CV2025
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
Nyle Siddiqui, Rohit Gupta, Sirnam Swetha +1
State space models (SSMs) have emerged as a competitive alternative to transformers in various tasks. Their linear complexity and hidden-state recurrence make them particularly att…
cs.CV2024
DLCR: A Generative Data Expansion Framework via Diffusion for Clothes-Changing Person Re-ID
Nyle Siddiqui, Florinel Alin Croitoru, Gaurav Kumar Nayak +2
With the recent exhibited strength of generative diffusion models, an open research question is if images generated by these models can be used to learn better visual representatio…