2 papers
cs.CV2026
Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images
Zheng Jiang, Yiming Chen, Nan He +4
Recent multimodal large language models (MLLMs) support Thinking with Images, invoking visual tools such as zooming and cropping to inspect image regions during inference. Yet thes…
cs.CV2026
MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning
Zheng Jiang, Heng Guo, Chengyu Fang +4
Medical Vision-Language Models (VLMs) hold immense promise for complex clinical tasks, but their reasoning capabilities are often constrained by text-only paradigms that fail to gr…