10 papers
PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models
Davood Soleymanzadeh, Kaidi Zhang, Zhiyuan Zhang +4
Vision-language-action models (VLAs) have shown strong promise for general-purpose robotic manipulation by mapping language instructions and vision observations directly to actions…
Imagining the Sense of Touch: Touch-Informed Manipulation via Imagined Tactile Representations
Zhiyuan Zhang, Adeesh Desai, Jyun-Chi Hu +7
Tactile sensing can substantially improve contact-rich robotic manipulation, yet its practical deployment remains limited by the fragility, calibration requirements, and maintenanc…
DREAM-Chunk: Reactive Action Chunking with Latent World Model
Wenxi Chen, Kaidi Zhang, Chi Lin +6
Action chunking has become a common interface for vision-language-action (VLA) models, enabling low-frequency policy inference to drive high-frequency robot execution. However, onc…
ReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous Driving
Zhiyuan Zhang, Yanlun Peng, Jianing Zhang +7
Reactive capability is a key property of data-driven behavior world model simulators for autonomous driving simulation systems. With this capability, simulated world agents can res…
ContactWorld: What Representations Matter for Vision-Tactile Latent World Models in Contact-Rich Manipulation
Zhiyuan Zhang, Pokuang Zhou, Kaidi Zhang +7
Contact-rich manipulation poses a fundamental challenge for world models: visual and tactile observations capture different aspects of physical interaction, and their utility depen…
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
Kaidi Zhang, Heng Zhang, Zhengtong Xu +7
Vision-Language-Action (VLA) models have demonstrated significant advantages in robotic manipulation. However, their reliance on vision and language often leads to suboptimal perfo…