254 citations · 362 across the 10 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
stat.ML2023
Nash Learning from Human Feedback
Rémi Munos, Michal Valko, Daniele Calandriello +14
Reinforcement learning from human feedback (RLHF) has emerged as the main paradigm for aligning large language models (LLMs) with human preferences. Typically, RLHF involves the in…
physics.plasm-ph2023
Towards practical reinforcement learning for tokamak magnetic control
Brendan D. Tracey, Andrea Michi, Yuri Chervonyi +15
Reinforcement learning (RL) has shown promising results for real-time control systems, including the domain of plasma magnetic control. However, there are still significant drawbac…