2 papers
cs.RO2026
CommCP: Efficient Multi-Agent Coordination via LLM-Based Communication with Conformal Prediction
Xiaopan Zhang, Zejin Wang, Zhixu Li +2
To complete assignments provided by humans in natural language, robots must interpret commands, generate and answer relevant questions for scene understanding, and manipulate targe…
cs.LG2026
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
Chenghua Huang, Lu Wang, Fangkai Yang +6
In this paper, we explore how directly pretraining a value model simplifies and stabilizes reinforcement learning from human feedback (RLHF). In reinforcement learning, value estim…