8 papers
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
David G. Shatwell, Sirnam Swetha, Mubarak Shah
Many real-world applications in digital forensics, urban monitoring, and environmental analysis require jointly reasoning about visual appearance, location, and time. Beyond standa…
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
Nyle Siddiqui, Rohit Gupta, Sirnam Swetha +1
State space models (SSMs) have emerged as a competitive alternative to transformers in various tasks. Their linear complexity and hidden-state recurrence make them particularly att…
Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
Younggun Kim, Sirnam Swetha, Fazil Kagdi +1
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks. However, these models often infer and reveal sensitive biometric attrib…
Cross-View Open-Vocabulary Object Detection in Aerial Imagery
Jyoti Kini, Rohit Gupta, Mubarak Shah
Traditional object detection models are typically trained on a fixed set of classes, limiting their flexibility and making it costly to incorporate new categories. Open-vocabulary…
GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
David G. Shatwell, Ishan Rajendrakumar Dave, Sirnam Swetha +1
Timestamp prediction aims to determine when an image was captured using only visual information, supporting applications such as metadata correction, retrieval, and digital forensi…
VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment
Shaina Raza, Ashmal Vayani, Aditya Jain +8
Detecting disinformation that blends manipulated text and images has become increasingly challenging, as AI tools make synthetic content easy to generate and disseminate. While mos…