activity
20242026
collaborators

9 papers

cs.CV2026

What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom

Yan Ma, Weiyu Zhang, Tianle Li +3

Vision tool-use reinforcement learning (RL) can equip vision language models with visual operators such as crop-and-zoom and achieves strong performance gains, yet it remains uncle…

cs.CV2026

One RL to See Them All: Visual Triple Unified Reinforcement Learning

Yan Ma, Linge Du, Xuyang Shen +7

Reinforcement learning (RL) is becoming an important direction for post-training vision-language models (VLMs), but public training methodologies for unified multimodal RL remain m…

cs.CV2025

Visual Programmability: A Guide for Code-as-Thought in Chart Understanding

Bohao Tang, Yan Ma, Fei Zhang +6

Chart understanding presents a critical test to the reasoning capabilities of Vision-Language Models (VLMs). Prior approaches face critical limitations: some rely on external tools…

cs.CL2025

Interaction as Intelligence: Deep Research With Human-AI Partnership

Lyumanshan Ye, Xiaojie Cai, Xinkai Wang +23

This paper introduces "Interaction as Intelligence" research series, presenting a reconceptualization of human-AI relationships in deep research tasks. Traditional approaches treat…

cs.CV2025

Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Zhaochen Su, Peng Xia, Hangyu Guo +12

Recent progress in multimodal reasoning has been significantly advanced by textual Chain-of-Thought (CoT), a paradigm where models conduct reasoning within language. This text-cent…

cs.CV2025

Thinking with Generated Images

Ethan Chern, Zhulin Hu, Steffi Chern +5

We present Thinking with Generated Images, a novel paradigm that fundamentally transforms how large multimodal models (LMMs) engage with visual reasoning by enabling them to native…