3 papers
cs.RO2026
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
Jinhui Ye, Fangjing Wang, Ning Gao +9
Large vision-language models (VLMs) excel at multimodal understanding but fall short when extended to embodied tasks, where instructions must be transformed into low-level motor ac…
cs.RO2025
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
Xinyi Chen, Yilun Chen, Yanwei Fu +26
We introduce InternVLA-M1, a unified framework for spatial grounding and robot control that advances instruction-following robots toward scalable, general-purpose intelligence. Its…
cs.CV2025
HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision
Shengli Zhou, Jianuo Zhu, Qilin Huang +3
3D Visual Question-Answering (3D VQA) is pivotal for models to perceive the physical world and perform spatial reasoning. Answer-centric supervision is a commonly used training met…