3 papers
cs.AI2026
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou +11
Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against dia…
cs.AI2026
IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations
Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar +10
Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision s…
cs.CV2025
Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
Parth Thakkar, Ankush Agarwal, Prasad Kasu +2
While Multi-modal Large Language Models (MLLMs) have shown impressive capabilities in document understanding tasks, their ability to locate and reason about fine-grained details wi…