Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning
Enhan Zhao, Wei Wu, Yuanrui Zhang +2
Spatial reasoning remains a persistent challenge for multimodal large language models (MLLMs). Existing approaches largely rely on large-scale, statically curated datasets, where a…
cs.CV2026
Understanding vs. Generation: Navigating Optimization Dilemma in Multimodal Models
Sen Ye, Mengde Xu, Shuyang Gu +3
Current research in multimodal models faces a key challenge where enhancing generative capabilities often comes at the expense of understanding, and vice versa. We analyzed this tr…