3 citations · 3 across the 10 of their papers we have counts for
1 paper · 2 filters
Yitong Chen, Shiduo Zhang, Jingjing Gong +1
Generating diverse images from sparse text is hard; generating compact actions from rich observations is easier. From the condition-target view, Vision-Language-Action (VLA) thus a…