5 papers
StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
Zhe Liu, Jinghua Hou, Yuxiang Lu +7
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting…
DreamAvoid: Critical-Phase Test-Time Dreaming to Avoid Failures in VLA Policies
Xianzhe Fan, Yuxiang Lu, Shenyuan Gao +4
Vision-Language-Action (VLA) models are often brittle in fine-grained manipulation, where minor action errors during the critical phases can rapidly escalate into irrecoverable fai…
FASTER: Rethinking Real-Time Flow VLAs
Yuxiang Lu, Zhe Liu, Xianzhe Fan +5
Real-time execution is crucial for deploying Vision-Language-Action (VLA) models in the physical world. Existing asynchronous inference methods primarily optimize trajectory smooth…
Utonia: Toward One Encoder for All Point Clouds
Yujia Zhang, Xiaoyang Wu, Yunhan Yang +6
We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward…
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
Xianzhe Fan, Shengliang Deng, Xiaoyang Wu +7
Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. How can we incorporate 3D informa…