5 papers
NativeMEM: Native Memory Compression for Long-Horizon Robotic Manipulation
Ziye Wang, Modi Shi, Chaojun Ni +5
How can pretrained Vision-Language-Action (VLA) models retain long-horizon visual histories with high-frequency updates without sacrificing efficiency? Existing approaches rely on…
StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning
Xiaowen Sun, Matthias Kerzel, Mengdi Li +3
Vision-language models (VLMs) have shown remarkable performance in various robotic tasks, as they can perceive visual information and understand natural language instructions. Howe…
Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
Jiaye Lin, Mengdi Li, Xufeng Zhao +4
Reward models trained through Reinforcement Learning from AI Feedback (RLAIF) methods frequently suffer from limited generalizability, which hinders the alignment performance of po…
HoloBrain-0 Technical Report
Xuewu Lin, Tianwei Lin, Yun Du +12
In this work, we introduce HoloBrain-0, a comprehensive Vision-Language-Action (VLA) framework that bridges the gap between foundation model research and reliable real-world robot…
Large Language Models for Orchestrating Bimanual Robots
Kun Chu, Xufeng Zhao, Cornelius Weber +3
Although there has been rapid progress in endowing robots with the ability to solve complex manipulation tasks, generating control policies for bimanual robots to solve tasks invol…