2 papers
cs.LG2024
Averaging log-likelihoods in direct alignment
Nathan Grinsztajn, Yannis Flet-Berliac, Mohammad Gheshlaghi Azar +8
To better align Large Language Models (LLMs) with human judgment, Reinforcement Learning from Human Feedback (RLHF) learns a reward model and then optimizes it using regularized RL…
cs.LG2022
Meta-learning from Learning Curves Challenge: Lessons learned from the First Round and Design of the Second Round
Manh Hung Nguyen, Lisheng Sun, Nathan Grinsztajn +1
Meta-learning from learning curves is an important yet often neglected research area in the Machine Learning community. We introduce a series of Reinforcement Learning-based meta-l…