3 papers
cs.RO2026
GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models
Yizhi Chen, Zhanxiang Cao, Xinyi Peng +14
Current Vision--Language--Action (VLA) models often optimize for semantic grounding, whereas executable manipulation requires geometry-aware spatial alignment and dynamic affordanc…
cs.CV2026
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
Yiheng Li, Zhuo Li, Ruibing Hou +4
Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed mo…
cs.CV2025
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
Yiheng Li, Ruibing Hou, Hong Chang +2
Human pose plays a crucial role in the digital age. While recent works have achieved impressive progress in understanding and generating human poses, they often support only a sing…