2 papers
cs.LG2023
Reinforcement Learning Fine-tuning of Language Models is Biased Towards More Extractable Features
Diogo Cruz, Edoardo Pona, Alex Holness-Tofts +4
Many capable large language models (LLMs) are developed via self-supervised pre-training followed by a reinforcement-learning fine-tuning phase, often based on human or AI feedback…
cs.LG2023★ 1 cited
Goodhart's Law in Reinforcement Learning
Jacek Karwowski, Oliver Hayman, Xingjian Bai +3
Implementing a reward function that perfectly captures a complex task in the real world is impractical. As a result, it is often appropriate to think of the reward function as a pr…