5 papers
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou +11
Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against dia…
IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations
Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar +10
Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision s…
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
Muntasir Wahed, Xiaona Zhou, Kiet A. Nguyen +5
Recent advancements in Large Language Models (LLMs) have significantly enhanced their code generation capabilities. However, their robustness against adversarial misuse, particular…
DocCHA: Towards LLM-Augmented Interactive Online diagnosis System
Xinyi Liu, Dachun Sun, Yi R. Fung +2
Despite the impressive capabilities of Large Language Models (LLMs), existing Conversational Health Agents (CHAs) remain static and brittle, incapable of adaptive multi-turn reason…
Uncovering Cross-Domain Recommendation Ability of Large Language Models
Xinyi Liu, Ruijie Wang, Dachun Sun +2
Cross-Domain Recommendation (CDR) seeks to enhance item retrieval in low-resource domains by transferring knowledge from high-resource domains. While recent advancements in Large L…