activity
20242026
collaborators

8 papers

cs.CY2026

Agentic AI Enhances Physician Trust in Clinical Decision Making

Zhiling Yan, Zhe Fang, David J King +10

Medical AI has shifted from reasoning to agentic AI, a new paradigm that autonomously invokes external tools during reasoning, rendering intermediate reasoning steps and tool outpu…

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…

cs.AI2026

OpenSkill: Open-World Self-Evolution for LLM Agents

Zhiling Yan, Dingjie Song, Hanrong Zhang +8

Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful trajectories, or verifier signa…

cs.CV2026

SegMix:Shuffle-based Feedback Learning for Semantic Segmentation of Pathology Images

Zhiling Yan, Sicheng Chen, Tianyi Zhang +3

Segmentation is a critical task in computational pathology, as it identifies areas affected by disease or abnormal growth and is essential for diagnosis and treatment. However, acq…

cs.AI2026

Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math

Dingjie Song, Tianlong Xu, Yi-Fan Zhang +6

Assessing student handwritten scratchwork is crucial for personalized educational feedback but presents unique challenges due to diverse handwriting, complex layouts, and varied pr…

cs.AI2026

LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric Evaluation

Zhiling Yan, Dingjie Song, Zhe Fang +4

The deployment of Large Language Models (LLMs) in high-stakes clinical settings demands rigorous and reliable evaluation. However, existing medical benchmarks remain static, suffer…