6 papers · 1 filter
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
Siqi Yang, Zilve Gao, Haibo Qiu +5
Multimodal Large Language Models (MLLMs) demonstrate significant potential but remain brittle in complex, long-chain visual reasoning tasks. A critical failure mode is "visual forg…
Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
Chi Zhang, Haibo Qiu, Qiming Zhang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Large Language Models (LLMs) and is now being applied to Vision-Langu…
Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
Xiaoxiao Ma, Feng Zhao, Pengyang Ling +6
In this work, we first revisit the sampling issues in current autoregressive (AR) image generation models and identify that image tokens, unlike text tokens, exhibit lower informat…
Counting Hallucinations in Diffusion Models
Shuai Fu, Jian Zhou, Qi Chen +7
Diffusion probabilistic models (DPMs) have demonstrated remarkable progress in generative tasks, such as image and video synthesis. However, they still often produce hallucinated s…
STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
Xiaoxiao Ma, Haibo Qiu, Guohui Zhang +4
Reinforcement learning has recently been explored to improve text-to-image generation, yet applying existing GRPO algorithms to autoregressive (AR) image models remains challenging…
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
Chi Zhang, Haibo Qiu, Qiming Zhang +3
The "thinking with images" paradigm represents a pivotal shift in the reasoning of Vision Language Models (VLMs), moving from text-dominant chain-of-thought to image-interactive re…