papers

Publications (9)

cs.CV2025

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs

Yan Shu, Chi Liu, Robin Chen +2

Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in various general-domain scenarios, such as visual question answering and image captioning. Rec…

cs.CV2026

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA

Mengzhuo Chen, Yan Shu, Chi Liu +4

We study whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both input types are aligned through a common reasoning interface. W…

cs.LG2025

Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning

Chi Liu, Derek Li, Yan Shu +4

While large language models show promise in medical applications, achieving expert-level clinical reasoning remains challenging due to the need for both accurate answers and transp…

cs.LG2025

Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling

Derek Li, Jiaming Zhou, Leo Maxime Brunswic +8

The pursuit of general-purpose artificial intelligence depends on large language models (LLMs) that can handle both structured reasoning and open-ended generation. We present Omni-…

cs.AI2026

Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher

Hongming Piao, Chi Liu, Mengzhuo Chen +5

Deep research and agent evolution serve as de-facto tasks for AI agents in real-world applications toward artificial general intelligence. The former enables autonomous retrieval a…

cs.LG2026

Best-of-Evidence: Best-of-N Selection under Partial Verification

Cenwei Zhang, Teng Fang, Yuxia Wang +3

BoN improves model outputs by sampling several candidates and selecting one with a proxy score, but it assumes that complete candidates can be evaluated reliably. Many vision-langu…

cs.DC2026

Talaria: Session-Aware Serverless Serving of Hundred-Billion-Parameter LLMs

Utopia Meng, Unicornt Zhao, Derek Li +2

Serverless multi-model LLM systems multiplex popularity-skewed model catalogs over shared GPU pools, yet typically schedule each request independently. Tool-using agents break this…

cs.CL2026

Hierarchical Chain-of-Thought: Enhancing LLM Reasoning Performance and Efficiency

Xingshuai Huang, Derek Li, Bahareh Nikpour +1

Chain-of-Thought (CoT) prompting has significantly improved the reasoning capabilities of large language models (LLMs). However, conventional CoT often relies on unstructured, flat…

cs.AI2025

Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs

Mohammad Ali Alomrani, Yingxue Zhang, Derek Li +14

Large language models (LLMs) have rapidly progressed into general-purpose agents capable of solving a broad spectrum of tasks. However, current models remain inefficient at reasoni…