collaborators

8 papers

cs.LG2026

GAC: Stabilizing Asynchronous RL Training for LLMs via Gradient Alignment Control

Haofeng Xu, Junwei Su, Yukun Tian +3

Asynchronous execution is essential for scaling reinforcement learning (RL) to modern large model workloads, including large language models and AI agents, but it can fundamentally…

cs.AI2026

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning

Zhe Qian, Nianbing Su, Zhonghua Wang +6

Current multimodal models often suffer from shallow reasoning, leading to errors caused by incomplete or inconsistent thought processes. To address this limitation, we propose Self…

cs.AI2026

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models

Zhe Qian, Yanbiao Ma, Zhuohan Ouyang +7

Multimodal Large Reasoning Models (MLRMs) have achieved remarkable strides in visual reasoning through test time compute scaling, yet long chain reasoning remains prone to hallucin…

cs.DC2026

Accelerating Compound LLM Training Workloads with Maestro

Xiulong Yuan, Hongqing Chen, Jiaxuan Peng +16

Compound LLM training workloads-such as knowledge distillation and multimodal LLM (MLLM) training-are gaining prominence. These typically comprise heterogeneous components differin…

cs.LG2026

Reasoning emerges from constrained inference manifolds in large language models

Yanbiao Ma, Fei Luo, Linfeng Zhang +10

Reasoning in large language models is predominantly evaluated through labeled benchmarks, conflating task performance with the quality of internal inference. Here we study reasonin…

cs.CV2026

Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding

Zhongxing Xu, Zhonghua Wang, Zhe Qian +10

Recent advancements in multimodal large reasoning models (MLRMs) have significantly improved performance in visual question answering. However, we observe that transition words (e.…