2 citations · 2 across the 1 of their papers we have counts for
1 paper
Leo McKee-Reid, Christoph Sträter, Maria Angelica Martinez +2
Previous work has shown that training "helpful-only" LLMs with reinforcement learning on a curriculum of gameable environments can lead models to generalize to egregious specificat…