7 papers · 1 filter
Counterfactually Safe Reinforcement Learning
Jingyi Li, Peng Wu, Chengchun Shi
Reinforcement learning algorithms are generally designed to maximize the expected return across a population. However, a policy that is optimal on average may be suboptimal for cer…
Perturbation is All You Need for Extrapolating Language Models
Zetai Cen, Jin Zhu, Xinwei Shen +1
This paper develops a statistical theory of extrapolation for large language models, by reinterpreting them through pre-post-additive noise models. In contrast to the standard auto…
Reinforcement Learning from Human Feedback: A Statistical Perspective
Pangpang Liu, Chengchun Shi, Will Wei Sun
Reinforcement learning from human feedback (RLHF) has emerged as a central framework for aligning large language models (LLMs) with human preferences. Despite its practical success…
Robust Reinforcement Learning from Human Feedback for Large Language Models Fine-Tuning
Kai Ye, Hongyi Zhou, Jin Zhu +2
Reinforcement learning from human feedback (RLHF) has emerged as a key technique for aligning the output of large language models (LLMs) with human preferences. To learn the reward…
Statistical Inference in Reinforcement Learning: A Selective Survey
Chengchun Shi
Reinforcement learning (RL) is concerned with how intelligence agents take actions in a given environment to maximize the cumulative reward they receive. In healthcare, applying RL…
Dual Active Learning for Reinforcement Learning from Human Feedback
Pangpang Liu, Chengchun Shi, Will Wei Sun
Aligning large language models (LLMs) with human preferences is critical to recent advances in generative artificial intelligence. Reinforcement learning from human feedback (RLHF)…