Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
Action with Visual Primitives
Weilong Guo, Yuchen Wang, Renping Zhou +5
Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robotic manipulation. A common design in current architectures maps language instructions an…
cs.RO2026
MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models
Hao Shi, Weiye Li, Bin Xie +6
Temporal modeling is essential for robotic manipulation, as effective control requires both memory of past interactions and imagination of future states. However, most VLA models r…