297 citations · 299 across the 5 of their papers we have counts for
1 paper · 2 filters
Nicolas Le Roux, Marc G. Bellemare, Jonathan Lebensold +7
We propose a new algorithm for fine-tuning large language models using reinforcement learning. Tapered Off-Policy REINFORCE (TOPR) uses an asymmetric, tapered variant of importance…