9 papers
Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments
Mingyu Liu, Zeju Li, Jiuhe Shu +4
The paper shows that smooth robot demonstrations can miss critical alignment moments, and proposes slowing down and resampling key motion segments, plus a spatio‑temporal feature c…
MARBLE: Multi-Aspect Reward Balance for Diffusion RL
Canyu Zhao, Hao Chen, Yunze Tong +3
Reinforcement learning fine-tuning has become the dominant approach for aligning diffusion models with human preferences. However, assessing images is intrinsically a multi-dimensi…
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
Hao Chen, Jiaming Liu, Zhonghao Yan +10
Robotic foundation models require reasoning over complex visual scenes to execute adaptive actions in dynamic environments. While recent studies on latent-reasoning Vision-Language…
AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking
Kui Wu, Hao Chen, Jinzhu Han +6
Realizing active visual tracking with a single unified model across diverse robots is challenging, as the physical constraints and motion dynamics vary drastically from one platfor…
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Yicheng Zou, Dongsheng Zhu, Lin Zhu +174
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancem…
AoE: Always-on Egocentric Human Video Collection for Embodied AI
Bowen Yang, Zishuo Li, Yang Sun +15
Embodied foundation models require large-scale, high-quality real-world interaction data for pre-training and scaling. However, existing data collection methods suffer from high in…