4 citations · 7 across the 6 of their papers we have counts for
1 paper · 1 filter
Jacek Karwowski, Younesse Kaddar, Zihuiwen Ye +2
When language models are trained by reinforcement learning (RL) to write probabilistic programs, they can artificially inflate their marginal-likelihood reward by producing program…