10 papers · 1 filter
DriftingVLA: Native One-Step Vision-Language-Action Generation via Per-Dimension Temporal Drifting
Yuxuan Gao, Shiqi Zhang, Yedong Shen +6
Conventional flow-based vision-language-action (VLA) models support expressive continuous action generation but rely on multi-step refinement to produce each action chunk, increasi…
TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction
Lei Jin, Yiding Ma, Xin Zhang +3
World Action Models (WAMs) combine future-state prediction with robot action generation, but existing approaches largely rely on visual futures. Visual prediction captures scene st…
ReTouch: Empowering Contact-Rich Dexterous Manipulation with Online-Refined Tactile Prediction
Shiqi Zhang, Xin Zhang, Yedong Shen +9
Fusing tactile signals has proven effective for contact-rich manipulation, enabling robots to perceive contact states and adapt to rapidly changing physical interactions. Yet effec…
WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory
Haisheng Su, Zongdai Liu, Xin Jin +13
World Action Models (WAMs) offer a promising paradigm for robotic manipulation by jointly modeling visual state transitions and robot actions. However, existing WAMs are constraine…
Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
Jianjie Fang, Yongyan Xu, Ziyou Wang +13
World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulators in which agents can perceive, act, fo…
GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation
Yuan Zhang, Shiqi Zhang, Yedong Shen +11
Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot emb…