2 papers
cs.CV2026
SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
Cheng Yin, Wang Xu, Junpeng Yang +8
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms…
cs.CV2026
Thinker: A vision-language foundation model for embodied intelligence
Baiyu Pan, Daqin Luo, Junpeng Yang +4
When large vision-language models are applied to the field of robotics, they encounter problems that are simple for humans yet error-prone for models. Such issues include confusion…