6 papers
Controlling Embedding Spaces with Text-Conditioned Transformations
Joseph Fioresi, Fabian Caba Heilbron, Pankaj Nathani +2
Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification. These embeddings compre…
Learning to Share: Selective Memory for Efficient Parallel Agentic Systems
Joseph Fioresi, Parth Parag Kulkarni, Ashmal Vayani +2
Agentic systems solve complex tasks by coordinating multiple agents that iteratively reason, invoke tools, and exchange intermediate results. To improve robustness and solution qua…
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
Sirnam Swetha, Rohit Gupta, Parth Parag Kulkarni +5
Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly…
MedRoute: RL-Based Dynamic Specialist Routing in Multi-Agent Medical Diagnosis
Ashmal Vayani, Parth Parag Kulkarni, Joseph Fioresi +2
Medical diagnosis using Large Multimodal Models (LMMs) has gained increasing attention due to capability of these models in providing precise diagnoses. These models generally comb…
GAEA: A Geolocation Aware Conversational Assistant
Ron Campos, Ashmal Vayani, Parth Parag Kulkarni +4
Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the us…
CityGuessr: City-Level Video Geo-Localization on a Global Scale
Parth Parag Kulkarni, Gaurav Kumar Nayak, Mubarak Shah
Video geolocalization is a crucial problem in current times. Given just a video, ascertaining where it was captured from can have a plethora of advantages. The problem of worldwide…