1 paper
Zihao Zheng, Sicheng Tian, Zhihao Mao +8
Vision-Language-Action (VLA) models have emerged as the mainstream of embodied intelligence. Recent VLA models have expanded their input modalities from 2D-only to 2D+3D paradigms,…