7 papers
Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks
Yang Liu, Kejia Zhang, Bingjie Yan +11
Large language models (LMs) offer broad generalization capabilities but require vast amounts of data and computational resources for domain-specific tasks; small models (SMs), in c…
EyeWorld: A Generative World Model of Ocular State and Dynamics
Ziyu Gao, Xinyuan Wu, Xiaolan Chen +10
Ophthalmic decision-making depends on subtle lesion-scale cues interpreted across multimodal imaging and over time, yet most medical foundation models remain static and degrade und…
EyeAgent: An Agentic AI System for Multimodal Clinical Decision Support in Ophthalmology
Danli Shi, Xiaolan Chen, Bingjie Yan +24
Artificial intelligence has shown promise in medical imaging, yet most existing systems lack flexibility, interpretability, and adaptability - challenges especially pronounced in o…
Empowering Locally Deployable Medical Agent via State Enhanced Logical Skills for FHIR-based Clinical Tasks
Wanrong Yang, Zhengliang Liu, Yuan Li +6
While Large Language Models demonstrate immense potential as proactive Medical Agents, their real-world deployment is severely bottlenecked by data scarcity under privacy constrain…
MInCo: Mitigating Information Conflicts in Distracted Visual Model-based Reinforcement Learning
Shiguang Sun, Hanbo Zhang, Zeyang Liu +4
Existing visual model-based reinforcement learning (MBRL) algorithms with observation reconstruction often suffer from information conflicts, making it difficult to learn compact r…
Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat
Pusheng Xu, Xia Gong, Xiaolan Chen +7
Purpose: To develop a bilingual multimodal visual question answering (VQA) benchmark for evaluating VLMs in ophthalmology. Methods: Ophthalmic image posts and associated captions p…