18 papers
SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment
Rong Xue, Jiageng Mao, Mingtong Zhang +1
Developing efficient and accurate visuomotor policies poses a central challenge in robotic imitation learning. While recent rectified flow approaches have advanced visuomotor polic…
Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
Zhenyu Zhao, Hongyi Jing, Xiawei Liu +7
From loco-motion to dextrous manipulation, humanoid robots have made remarkable strides in demonstrating complex full-body capabilities. However, the majority of current robot lear…
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models
Yanru Wu, Weiduo Yuan, Ang Qi +3
Reinforcement Learning (RL) has shown great potential in refining robotic manipulation policies, yet its efficacy remains strongly bottlenecked by the difficulty of designing gener…
DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models
Emily Yue-Ting Jia, Weiduo Yuan, Tianheng Shi +3
Robotic manipulation requires sophisticated commonsense reasoning, a capability naturally possessed by large-scale Vision-Language Models (VLMs). While VLMs show promise as zero-sh…
: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation
Songlin Wei, Hongyi Jing, Boqian Li +12
We introduce (Psi-Zero), an open foundation model to address challenging humanoid loco-manipulation tasks. While existing approaches often attempt to address this fundamenta…
InstantSfM: Towards GPU-Native SfM for the Deep Learning Era
Jiankun Zhong, Zitong Zhan, Quankai Gao +6
Structure-from-Motion (SfM) is a fundamental technique for recovering camera poses and scene structure from multi-view imagery, serving as a critical upstream component for applica…