4 papers
Medical thinking with multiple images
Zonghai Yao, Benlu Wang, Yifan Zhang +8
Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a…
Efficient Test-Time Scaling via Temporal Reasoning Aggregation
Jiakun Li, Xingwei He, Kefan Li +3
Test-time scaling improves the reasoning performance of large language models but often results in token-inefficient overthinking, where models continue reasoning beyond what is ne…
Rethinking Patient Education as Multi-turn Multi-modal Interaction
Zonghai Yao, Zhipeng Tang, Chengtao Lin +5
Most medical multimodal benchmarks focus on static tasks such as image question answering, report generation, and plain-language rewriting. Patient education is more demanding: sys…
A Psychology-based Unified Dynamic Framework for Curriculum Learning
Guangyu Meng, Qingkai Zeng, John P. Lalor +1
Directly learning from examples of varying difficulty levels is often challenging for both humans and machine learning models. A more effective strategy involves exposing learners…