4 papers
Multi-Rollout On-Policy Distillation via Peer Successes and Failures
Weichen Yu, Xiaomin Li, Yizhou Zhao +8
Large language models are often post-trained with sparse verifier rewards, which indicate whether a sampled trajectory succeeds but provide limited guidance about where reasoning s…
CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
Ke Sun, Yizhou Zhao, Jiayi Xin +2
Context or prompt-level reweighting has emerged as a central algorithmic lever in Reinforcement Learning with Verified Rewards (RLVR) for improving the reasoning capability of larg…
ConsistNav: Closing the Action Consistency Gap in Zero-Shot Object Navigation with Semantic Executive Control
Haosen Wang, Zhenyang Li, Yinqiang Zhang +9
Zero-shot object navigation has advanced rapidly with open-vocabulary detectors, image--text models, and language-guided exploration. However, even after current methods detect a p…
RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models
Hao Wu, Yuqi Li, Yuan Gao +10
Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly aligned with the capabilities t…