2 papers
cs.CV2026
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
Jian Zhang, Shijie Zhou, Bangya Liu +2
Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This limitation arises from their ina…
cs.CV2025
Breaking the Box: Enhancing Remote Sensing Image Segmentation with Freehand Sketches
Ying Zang, Yuncan Gao, Jiangi Zhang +7
This work advances zero-shot interactive segmentation for remote sensing imagery through three key contributions. First, we propose a novel sketch-based prompting method, enabling…