3 papers
cs.LG2026
Information-Directed Offline-to-Online Reinforcement Learning
Keru Chen
Decision-making from offline datasets typically warm-starts a policy or score model from fixed offline data and then refines it with limited online interaction. Offline data reduce…
cs.LG2026
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
Keru Chen, Jun Luo, Sen Lin +4
Hierarchical Instruction Following (HIF) refers to the problem of prompting large language models with a priority-ordered stack of instructions. Standard methods like RLHF and DPO…
cs.LG2026
Towards Fast Safe Online Reinforcement Learning via Policy Finetuning
Keru Chen, Honghao Wei, Zhigang Deng +1
The high costs and risks involved in extensive environment interactions hinder the practical application of current online safe reinforcement learning (RL) methods. While offline s…