3 papers
cs.RO2026
Non-Markovian Long-Horizon Robot Manipulation via Keyframe Chaining
Yipeng Chen, Wentao Tan, Lei Zhu +4
Existing Vision-Language-Action (VLA) models often struggle to generalize to long-horizon tasks due to their heavy reliance on immediate observations. While recent studies incorpor…
cs.AI2025
BLM: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
Wentao Tan, Bowen Wang, Heng Zhi +15
Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs ge…
cs.RO2025
RoboBERT: An End-to-end Multimodal Robotic Manipulation Model
Sicheng Wang, Sheng Liu, Weiheng Wang +2
Embodied intelligence seamlessly integrates vision, language, and action.~However, most multimodal robotic models rely on massive fine-tuning, incurring high time and hardware cost…