most citedGeoChat: Grounded Large Vision-Language Model for Remote Sensing

6 citations · 6 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Counterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs

Yifan Lu, Adinath Dukre, Abhijit Das +4

Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically u…

cs.CV2026

Graph-of-Differences: Anatomy-Structured Difference Alignment for Medical Image Re-Identification

Nichula Wasalathilaka, Abhijit Das, Imran Razzak +1

Medical image re-identification (MedReID) enables longitudinal patient linkage but remains vulnerable to shortcut learning and often produces decisions that clinicians cannot audit…

cs.CV2026

PROTON: Prototype-Based Test-Time Online OOD Detection for Medical VLMs

Abhijit Das, Nichula Wasalathilaka, Yifan Lu +4

Medical vision-language models (VLMs) enable zero-shot clinical image classification, yet reliably detecting out-of-distribution (OOD) inputs at deployment remains an open problem.…

cs.CV2026

B-MoE: A Body-Part-Aware Mixture-of-Experts "All Parts Matter" Approach to Micro-Action Recognition

Nishit Poddar, Aglind Reka, Diana-Laura Borza +4

Micro-actions, fleeting and low-amplitude motions, such as glances, nods, or minor posture shifts, carry rich social meaning but remain difficult for current action recognition mod…

cs.CV2026

Narrating For You: Prompt-guided Audio-visual Narrating Face Generation Employing Multi-entangled Latent Space

Aashish Chandra, Aashutosh A, Abhijit Das

We present a novel approach for generating realistic speaking and talking faces by synthesizing a person's voice and facial movements from a static image, a voice profile, and a ta…

cs.CV20236 cited

GeoChat: Grounded Large Vision-Language Model for Remote Sensing

Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer +3

Recent advancements in Large Vision-Language Models (VLMs) have shown great promise in natural image domains, allowing users to hold a dialogue about given visual content. However,…