3 papers
cs.CV2026
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
Sirnam Swetha, Rohit Gupta, Parth Parag Kulkarni +5
Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly…
cs.CV2026
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
David G. Shatwell, Sirnam Swetha, Mubarak Shah
Many real-world applications in digital forensics, urban monitoring, and environmental analysis require jointly reasoning about visual appearance, location, and time. Beyond standa…
cs.CV2025
GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
David G. Shatwell, Ishan Rajendrakumar Dave, Sirnam Swetha +1
Timestamp prediction aims to determine when an image was captured using only visual information, supporting applications such as metadata correction, retrieval, and digital forensi…