353 citations · 631 across the 9 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2023
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Stephen Casper, Xander Davies, Claudia Shi +29
Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of…
cs.AI2023
Thinker: Learning to Plan and Act
Stephen Chung, Ivan Anokhin, David Krueger
We propose the Thinker algorithm, a novel approach that enables reinforcement learning agents to autonomously interact with and utilize a learned world model. The Thinker algorithm…