1 paper · 1 filter
Leo McKee-Reid, Christoph Sträter, Maria Angelica Martinez +2
Previous work has shown that training "helpful-only" LLMs with reinforcement learning on a curriculum of gameable environments can lead models to generalize to egregious specificat…