6 papers
PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models
Davood Soleymanzadeh, Kaidi Zhang, Zhiyuan Zhang +4
Vision-language-action models (VLAs) have shown strong promise for general-purpose robotic manipulation by mapping language instructions and vision observations directly to actions…
DREAM-Chunk: Reactive Action Chunking with Latent World Model
Wenxi Chen, Kaidi Zhang, Chi Lin +6
Action chunking has become a common interface for vision-language-action (VLA) models, enabling low-frequency policy inference to drive high-frequency robot execution. However, onc…
ContactWorld: What Representations Matter in Vision-Tactile World Models for Contact-Rich Manipulation
Zhiyuan Zhang, Pokuang Zhou, Kaidi Zhang +6
Contact-rich manipulation requires world models to capture complex interaction dynamics from heterogeneous visual and tactile observations, yet the representation properties that e…
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
Kaidi Zhang, Heng Zhang, Zhengtong Xu +7
Vision-Language-Action (VLA) models have demonstrated significant advantages in robotic manipulation. However, their reliance on vision and language often leads to suboptimal perfo…
PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation
Pengyuan Guo, Zhonghao Mai, Zhengtong Xu +8
Recent advances in vision-language models (VLMs) have enabled increasing progress in real-world robot manipulation. However, long-horizon manipulation in unstructured environments…
VibeCheck: Using Active Acoustic Tactile Sensing for Contact-Rich Manipulation
Kaidi Zhang, Do-Gon Kim, Eric T. Chang +6
The acoustic response of an object can reveal a lot about its global state, for example its material properties or the extrinsic contacts it is making with the world. In this work,…