2 papers
cs.CV2026
Test-Time Training for Robust Text-Guided Open-Vocabulary Object Counting
Hao-Yuan Ma, Yuda Zou, Li Zhang +1
Text-guided Open-vocabulary Object Counting (TOOC) enables counting arbitrary object categories specified by text prompts, offering substantially greater flexibility than conventio…
cs.CV2026
Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence
Yufei Zheng, Xuhan Zhu, Zide Liu +9
Spatial understanding of the physical world from 2D visual inputs hinges on two complementary forms of geometric knowledge: holistic 3D structural perception and fine-grained metri…