2 papers
cs.CV2026
SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
Cheng Yin, Wang Xu, Junpeng Yang +8
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms…
cs.LG2025
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
Cheng Yin, Yankai Lin, Wang Xu +4
Does Chain-of-Thought (CoT) reasoning genuinely improve Vision Language Action (VLA) models, or does it merely add overhead? Existing CoT-VLA systems report limited and inconsisten…