distributionally robust reinforcement learning 1exploration-exploitation tradeoff 1interactive data collection 1robust inventory control 1robust markov decision processes 1sample complexity 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms
Miao Lu, Han Zhong, Tong Zhang +1
The paper studies reinforcement learning where the learner must be robust to differences between training and deployment environments, using interactive data collection and proposi…
cs.LG2025
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
Heyang Zhao, Chenlu Ye, Quanquan Gu +1
Reverse-Kullback-Leibler (KL) regularization has emerged to be a predominant technique used to enhance policy optimization in reinforcement learning (RL) and reinforcement learning…
cs.LG2024
Online Iterative Reinforcement Learning from Human Feedback with General Preference Model
Chenlu Ye, Wei Xiong, Yuheng Zhang +3
We investigate Reinforcement Learning from Human Feedback (RLHF) in the context of a general preference oracle. In particular, we do not assume the existence of a reward function a…