Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling
Zelin Zhao, Min Shi, Bo Yuan +5
World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-Worl…
cs.CV2026
Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI
Lama Moukheiber, Caleb M. Yeung, Haotian Xue +3
Spatial reasoning and visual grounding are core capabilities for vision-language models (VLMs), yet most medical VLMs produce predictions without transparent reasoning or spatial e…