3 papers
cs.CV2026
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
Chi Zhang, Haibo Qiu, Qiming Zhang +3
The "thinking with images" paradigm represents a pivotal shift in the reasoning of Vision Language Models (VLMs), moving from text-dominant chain-of-thought to image-interactive re…
cs.CV2025
Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
Chi Zhang, Haibo Qiu, Qiming Zhang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Large Language Models (LLMs) and is now being applied to Vision-Langu…
cs.CV2025
InstructVEdit: A Holistic Approach for Instructional Video Editing
Chi Zhang, Chengjian Feng, Feng Yan +5
Video editing according to instructions is a highly challenging task due to the difficulty in collecting large-scale, high-quality edited video pair data. This scarcity not only li…