8 papers
Thinking with Drafting: Optical Decompression via Logical Reconstruction
Jingxuan Wei, Honghao He, Caijun Jia +9
Existing multimodal large language models have achieved high-fidelity visual perception and exploratory visual generation. However, a precision paradox persists in complex reasonin…
How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning
Xiangxiang Zhang, Caijun Jia, Siyuan Li +9
Solving complex geometric problems inherently requires interleaved reasoning: a tight alternation between constructing diagrams and performing logical deductions. Although recent M…
HarmonyCell: Automating Single-Cell Perturbation Modeling under Semantic and Distribution Shifts
Wenxuan Huang, Mingyu Tsoi, Yanhao Huang +10
Single-cell perturbation studies face dual heterogeneity bottlenecks: (i) semantic heterogeneity--identical biological concepts encoded under incompatible metadata schemas across d…
Canvas-of-Thought: Grounding Reasoning via Mutable Structured States
Lingzhuang Sun, Yuxia Zhu, Ruitong Liu +10
While Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), relying solely on linear text sequences re…
Guided Verifier: Collaborative Multimodal Reasoning via Dynamic Process Supervision
Lingzhuang Sun, Ruitong Liu, Yuxia Zhu +5
Reinforcement Learning (RL) has emerged as a pivotal mechanism for enhancing the complex reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevailing par…
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
Xiangxiang Zhang, Jingxuan Wei, Donghong Zhong +31
Existing Vision-Language Models often struggle with complex, multi-question reasoning tasks where partial correctness is crucial for effective learning. Traditional reward mechanis…