6 papers
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Shuailei Ma, Jiaqi Liao, Xinyang Wang +24
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inheren…
From Foundation to Application: Improving VLA Models in Practice
Wei Wu, Fangjing Wang, Fan Lu +21
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bri…
Vision Pretraining for Dense Spatial Perception
Zelin Fu, Bin Tan, Changjiang Sun +6
Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observat…
A Pragmatic VLA Foundation Model
Wei Wu, Fan Lu, Yunnan Wang +22
Offering great potential in robotic manipulation, a capable Vision-Language-Action (VLA) foundation model is expected to faithfully generalize across tasks and platforms while ensu…
Learning Consistent Taxonomic Classification through Hierarchical Reasoning
Zhenghong Li, Kecheng Zheng, Haibin Ling
While Vision-Language Models (VLMs) excel at visual understanding, they often fail to grasp hierarchical knowledge. This leads to common errors where VLMs misclassify coarser taxon…
The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents
Ziyu Wang, Chenyuan Liu, Yushun Xiang +16
Recently, with the rapid development of robot learning and imitation learning, numerous datasets and methods have emerged. However, these datasets and their task designs often lack…