activity
20242026
collaborators

12 papers

cs.CV2026

Medical thinking with multiple images

Zonghai Yao, Benlu Wang, Yifan Zhang +8

Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a…

cs.AI2026

ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes Care

Zonghai Yao, Talha Chafekar, Junda Wang +5

Real-world adoption of closed-loop insulin delivery systems (CLIDS) in type 1 diabetes remains low, driven not by technical failure, but by diverse behavioral, psychosocial, and so…

cs.IR2026

TARSE: Test-Time Adaptation via Retrieval of Skills and Experience for Reasoning Agents

Junda Wang, Zonghai Tao, Hansi Zeng +3

Complex clinical decision making often fails not because a model lacks facts, but because it cannot reliably select and apply the right procedural knowledge and the right prior exa…

cs.AI2026

ESTAR: Early-Stopping Token-Aware Reasoning For Efficient Inference

Junda Wang, Zhichao Yang, Dongxu Zhang +2

Large reasoning models (LRMs) achieve state-of-the-art performance by generating long chains-of-thought, but often waste computation on redundant reasoning after the correct answer…

cs.CL2026

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations

Benlu Wang, Iris Xia, Yifan Zhang +6

Large language models (LLMs) have demonstrated promising performance on medical benchmarks; however, their ability to perform medical calculations, a crucial aspect of clinical dec…

cs.AI2026

MedQA-CS: Objective Structured Clinical Examination (OSCE)-Style Benchmark for Evaluating LLM Clinical Skills

Zonghai Yao, Zihao Zhang, Chaolong Tang +8

Artificial intelligence (AI) and large language models (LLMs) in healthcare require advanced clinical skills (CS), yet current benchmarks fail to evaluate these comprehensively. We…