23 citations · 23 across the 2 of their papers we have counts for
3 papers · 1 filter
Weak-to-Strong Generalization is Nearly Inevitable (in Linear Models)
Scott Geng, Dutch Hansen, Jerry Li
Weak-to-strong generalization is a phenomenon in post-training whereby a strong student model, when finetuned solely with feedback from a weaker teacher, can not only surpass the t…
Controlling Commercial Cooling Systems Using Reinforcement Learning
Jerry Luo, Cosmin Paduraru, Octavian Voicu +33
This paper is a technical overview of DeepMind and Google's recent work on reinforcement learning for controlling commercial cooling systems. Building on expertise that began with…
An empirical investigation of the challenges of real-world reinforcement learning
Gabriel Dulac-Arnold, Nir Levine, Daniel J. Mankowitz +4
Reinforcement learning (RL) has proven its worth in a series of artificial domains, and is beginning to show some successes in real-world scenarios. However, much of the research a…