294 citations · 356 across the 26 of their papers we have counts for
Showing 2023 · cs.LGShow all
2 papers · 2 filters
cs.LG2023★ 1 cited
Offline Retraining for Online RL: Decoupled Policy Learning to Mitigate Exploration Bias
Max Sobol Mark, Archit Sharma, Fahim Tajwar +3
It is desirable for policies to optimistically explore new states and behaviors during online reinforcement learning (RL) or fine-tuning, especially when prior offline data does no…
cs.LG2023★ 294 cited
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell +3
While large-scale unsupervised language models (LMs) learn broad world knowledge and some reasoning skills, achieving precise control of their behavior is difficult due to the comp…