6 papers
Best-of-Evidence: Best-of-N Selection under Partial Verification
Cenwei Zhang, Teng Fang, Yuxia Wang +3
BoN improves model outputs by sampling several candidates and selecting one with a proxy score, but it assumes that complete candidates can be evaluated reliably. Many vision-langu…
Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher
Hongming Piao, Chi Liu, Mengzhuo Chen +5
Deep research and agent evolution serve as de-facto tasks for AI agents in real-world applications toward artificial general intelligence. The former enables autonomous retrieval a…
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents
Jia Deng, Yimeng Chen, Xiaoqing Xiang +9
Training deep search agents requires verifiable questions whose answers remain unavailable until sufficient evidence has been acquired through search. Existing synthesis methods of…
UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA
Mengzhuo Chen, Yan Shu, Chi Liu +4
We study whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both input types are aligned through a common reasoning interface. W…
Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
Yan Shu, Chi Liu, Robin Chen +2
Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in various general-domain scenarios, such as visual question answering and image captioning. Rec…
Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
Chi Liu, Derek Li, Yan Shu +4
While large language models show promise in medical applications, achieving expert-level clinical reasoning remains challenging due to the need for both accurate answers and transp…