5 papers · 1 filter
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
Xianzhe Fan, Shengliang Deng, Xiaoyang Wu +7
Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. How can we incorporate 3D informa…
AVGGT: Rethinking Global Attention for Accelerating VGGT
Xianbing Sun, Zhikai Zhu, Zhengyu Lou +5
Models such as VGGT and have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-at…
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
Wenyao Zhang, Hongsi Liu, Zekun Qi +11
Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot ma…
DexVLG: Dexterous Vision-Language-Grasp Model at Scale
Jiawei He, Danshi Li, Xinqiang Yu +7
As large models gain traction, vision-language-action (VLA) systems are enabling robots to tackle increasingly complex tasks. However, limited by the difficulty of data collection,…
Implicit and Explicit Language Guidance for Diffusion-based Visual Perception
Hefeng Wang, Jiale Cao, Jin Xie +2
Text-to-image diffusion models have shown powerful ability on conditional image synthesis. With large-scale vision-language pre-training, diffusion models are able to generate high…