12 papers
Latent Action as Intention Enables Efficient Future Imagination for World Action Models
Xiang Li, Yupeng Zheng, Songen Gu +11
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes t…
VT-WAM: Visual-Tactile World Action Model for Contact-Rich Manipulation
Shuai Tian, Yupeng Zheng, Yuhang Zheng +7
Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observat…
TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation
Yujie Zang, Yuhang Zheng, Xian Nie +7
Contact-rich manipulation requires robots to continuously perceive and regulate evolving physical interactions under dynamic contact transitions or complex surface geometries. Rece…
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
Yupeng Zheng, Xiang Li, Songen Gu +12
Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency and a lack of high-level know…
VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
Songen Gu, Yuhang Zheng, Weize Li +6
Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to…
Mimir: Hierarchical Goal-Driven Diffusion with Uncertainty Propagation for End-to-End Autonomous Driving
Zebin Xing, Yupeng Zheng, Qichao Zhang +5
End-to-end autonomous driving has emerged as a pivotal direction in the field of autonomous systems. Recent works have demonstrated impressive performance by incorporating high-lev…