4 papers · 1 filter
Self-supervised Hierarchical Visual Reasoning with World Model
Yuanfei Xu, Lin Liu, Wengang Zhou +2
3D open-world environments with adversarial opponents remain a core challenge for reinforcement learning due to their vast state spaces. Effective reasoning representations are ess…
SGA-MCTS: Decoupling Planning from Execution via Training-Free Atomic Experience Retrieval
Xin Xie, Dongyun Xue, Wuguannan Yao +5
LLM-powered systems require complex multi-step decision-making abilities to solve real-world tasks, yet current planning approaches face a trade-off between the high latency of inf…
Search-Based Credit Assignment for Offline Preference-Based Reinforcement Learning
Xiancheng Gao, Yufeng Shi, Wengang Zhou +1
Offline reinforcement learning refers to the process of learning policies from fixed datasets, without requiring additional environment interaction. However, it often relies on wel…
Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks
Ruopei Sun, Jianfeng Cai, Jinhua Zhu +5
RLHF has emerged as a predominant approach for aligning artificial intelligence systems with human preferences, demonstrating exceptional and measurable efficacy in instruction fol…