4 papers
ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies
Jianming Ma, Rongjun Jin, Xiaxi Si +3
Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate h…
GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models
Yizhi Chen, Zhanxiang Cao, Xinyi Peng +14
Current Vision--Language--Action (VLA) models often optimize for semantic grounding, whereas executable manipulation requires geometry-aware spatial alignment and dynamic affordanc…
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
Yiheng Li, Zhuo Li, Ruibing Hou +4
Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed mo…
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
Yiheng Li, Ruibing Hou, Hong Chang +2
Human pose plays a crucial role in the digital age. While recent works have achieved impressive progress in understanding and generating human poses, they often support only a sing…