2 papers
cs.CV2026
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
Shaoxiong Zhan, Yanlin Lai, Zheng Liu +6
Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability m…
cs.CV2026
Truth in the Few: High-Value Data Selection for Efficient Multi-Modal Reasoning
Shenshen Li, Xing Xu, Kaiyuan Deng +3
While multi-modal large language models (MLLMs) have made significant progress in complex reasoning tasks via reinforcement learning, it is commonly believed that extensive trainin…