activity
20242026
collaborators

6 papers

cs.CV2026

Controlling Embedding Spaces with Text-Conditioned Transformations

Joseph Fioresi, Fabian Caba Heilbron, Pankaj Nathani +2

Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification. These embeddings compre…

cs.MA2026

Learning to Share: Selective Memory for Efficient Parallel Agentic Systems

Joseph Fioresi, Parth Parag Kulkarni, Ashmal Vayani +2

Agentic systems solve complex tasks by coordinating multiple agents that iteratively reason, invoke tools, and exchange intermediate results. To improve robustness and solution qua…

cs.CV2026

VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues

Sirnam Swetha, Rohit Gupta, Parth Parag Kulkarni +5

Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly…

eess.IV2026

MedRoute: RL-Based Dynamic Specialist Routing in Multi-Agent Medical Diagnosis

Ashmal Vayani, Parth Parag Kulkarni, Joseph Fioresi +2

Medical diagnosis using Large Multimodal Models (LMMs) has gained increasing attention due to capability of these models in providing precise diagnoses. These models generally comb…

cs.CV2025

GAEA: A Geolocation Aware Conversational Assistant

Ron Campos, Ashmal Vayani, Parth Parag Kulkarni +4

Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the us…

cs.CV2024

CityGuessr: City-Level Video Geo-Localization on a Global Scale

Parth Parag Kulkarni, Gaurav Kumar Nayak, Mubarak Shah

Video geolocalization is a crucial problem in current times. Given just a video, ascertaining where it was captured from can have a plethora of advantages. The problem of worldwide…