1 citations · 1 across the 14 of their papers we have counts for
Showing cs.ROShow all
3 papers · 1 filter
cs.RO2026
What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency
Luoyang Sun, Guoyang Xia, Fengfa Li +9
Vision-Language-Action (VLA) models combine a pretrained vision encoder, a language backbone, and an action head, but their relative contribution has not been established under con…
cs.RO2026
VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction
Hongjin Ji, Guoyang Xia, Luoyang Sun +2
Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in cl…
cs.RO2025
Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
Xianchao Zeng, Xinyu Zhou, Youcheng Li +5
Vision-Language-Action (VLA) models have recently achieved remarkable progress in robotic manipulation, yet they remain limited in failure diagnosis and learning from failures. Add…