2 papers
cs.RO2026
GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models
Yizhi Chen, Zhanxiang Cao, Xinyi Peng +14
Current Vision--Language--Action (VLA) models often optimize for semantic grounding, whereas executable manipulation requires geometry-aware spatial alignment and dynamic affordanc…
cs.CV2024
1st Place Solution of Multiview Egocentric Hand Tracking Challenge ECCV2024
Minqiang Zou, Zhi Lv, Riqiang Jin +4
Multi-view egocentric hand tracking is a challenging task and plays a critical role in VR interaction. In this report, we present a method that uses multi-view input images and cam…