2 papers
cs.AI2026
Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning
Peng, Lee, Yin Zhang +9
Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs). However, directly applying RL to rollout multimodal…
cs.CV2026
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
Zishan Liu, Ruoxi Zang, Yanglin Zhang +5
Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle with spatial reasoning, ofte…