6 papers
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
Lianyu Hu, Shengqian Qin, Zeqin Liao +4
Chain-of-thought (CoT) reasoning has enabled multi-modal large language models (MLLMs) to tackle complex visual reasoning tasks by generating explicit intermediate reasoning steps…
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Shuimu Chen, Yuteng Chen, Yuanshen Guan +7
Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evide…
PhotoCraft: Agentic Reasoning with Hierarchical Self-Evolving Memory for Deep Image Search
Kailin Lyu, Zhiqiang Yuan, Jianwei He +9
Deep Image Search requires multi-step reasoning over rich contextual cues, such as time, location, and event relations. However, most existing LLM-based agents are stateless and re…
EHRWorld: A Patient-Centric Medical World Model for Long-Horizon Clinical Trajectories
Linjie Mu, Zhongzhen Huang, Yannian Gu +3
World models offer a principled framework for simulating future states under interventions, but realizing such models in complex, high-stakes domains like medicine remains challeng…
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration
Yakun Zhu, Yutong Huang, Shengqian Qin +3
Medical calculators are fundamental to quantitative, evidence-based clinical practice. However, their real-world use is an adaptive, multi-stage process, requiring proactive EHR da…
MMXU: A Multi-Modal and Multi-X-ray Understanding Dataset for Disease Progression
Linjie Mu, Zhongzhen Huang, Shengqian Qin +3
Large vision-language models (LVLMs) have shown great promise in medical applications, particularly in visual question answering (MedVQA) and diagnosis from medical images. However…