6 papers
LUT: Latent Utility Training for Visual Reasoning
Jiaxuan Kang, Siyu Chen, Mingda Li +6
Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reasoning methods introduce hidden…
Thinking in Video: Can Video Generators Really Reason About the Real World?
Yongheng Zhang, Guang Yang, Ruihan Hou +12
Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about…
HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization
Tao Cheng, Shi-Zhe Chen, Hao Zhang +3
Chain-of-Thought (CoT) reasoning significantly elevates the complex problem-solving capabilities of multimodal large language models (MLLMs). However, adapting CoT to vision typica…
Exploring Autonomous Agentic Data Engineering for Model Specialization
Yujie Luo, Xiangyuan Ru, Jingsheng Zheng +10
Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data…
PRISM-MCTS: Learning from Reasoning Trajectories with Metacognitive Reflection
Siyuan Cheng, Bozhong Tian, YanChao Hao +1
PRISM-MCTS: Learning from Reasoning Trajectories with Metacognitive Reflection Siyuan Cheng, Bozhong Tian, Yanchao Hao, Zheng Wei Published: 06 Apr 2026, Last Modified: 06 Apr 2026…
Efficient Document Parsing via Parallel Token Prediction
Lei Li, Ze Zhao, Meng Li +6
Document parsing, as a fundamental yet crucial vision task, is being revolutionized by vision-language models (VLMs). However, the autoregressive (AR) decoding inherent to VLMs cre…