2 papers
cs.RO2026
Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?
Guoheng Sun, Kaixi Feng, Shwai He +8
Vision-Language-Action (VLA) models enable instruction-driven robotic manipulation, but they inherit oversized language backbones from pretrained VLMs whose capacity far exceeds wh…
cs.CV2026
ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models
Guoheng Sun, Tingting Du, Kaixi Feng +6
Vision-Language-Action (VLA) models enable instruction-following robotic manipulation, but they are typically pretrained on 2D data and lack 3D spatial understanding. An effective…