56 citations · 61 across the 3 of their papers we have counts for
3 papers
Zephyr: Direct Distillation of LM Alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert +11
We aim to produce a smaller language model that is aligned to user intent. Previous research has shown that applying distilled supervised fine-tuning (dSFT) on larger models signif…
Graph augmented Deep Reinforcement Learning in the GameRLand3D environment
Edward Beeching, Maxim Peter, Philippe Marcotte +4
We address planning and navigation in challenging 3D video games featuring maps with disconnected regions reachable by agents using special actions. In this setting, classical symb…
Godot Reinforcement Learning Agents
Edward Beeching, Jilles Debangoye, Olivier Simonin +1
We present Godot Reinforcement Learning (RL) Agents, an open-source interface for developing environments and agents in the Godot Game Engine. The Godot RL Agents interface allows…