From the 1 of 5 linked papers with an AI index.
5 papers
UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation
Junjie Lu, Xinyao Qin, Yuhua Jiang +6
The paper introduces UniSteer, a framework that converts human corrective actions into noise targets to guide a lightweight noise-prediction actor while simultaneously training it…
Beyond Monotonic Progress: Retry-Supervised Value Learning for Robot Imitation
Xinyao Qin, Junjie Lu, Kaixin Wang +7
Human demonstrations for robot imitation learning often contain mistakes and corrective behaviors, such as imprecise grasps, object misalignment, unstable contact, and repeated att…
Reinforcing VLAs in Task-Agnostic World Models
Yucen Wang, Rui Yu, Fengming Zhang +5
Post-training Vision-Language-Action (VLA) models via reinforcement learning (RL) in learned world models has emerged as an effective strategy to adapt to new tasks without costly…
Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
Junjie Lu, Yuliang Liu, Chaofeng Qu +4
Current approaches for strengthening LLM reasoning tend to introduce a training bias toward human-like reasoning trajectories. In step-wise preference optimization, in particular,…
AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
Yuliang Liu, Junjie Lu, Zhaoling Chen +10
Current approaches for training Process Reward Models (PRMs) often involve breaking down responses into multiple reasoning steps using rule-based techniques, such as using predefin…