18 citations · 23 across the 5 of their papers we have counts for
3 papers · 1 filter
Mastering Board Games by External and Internal Planning with Language Models
John Schultz, Jakub Adamek, Matej Jusup +13
Advancing planning and reasoning capabilities of Large Language Models (LLMs) is one of the key prerequisites towards unlocking their potential for performing reliably in complex a…
West-of-N: Synthetic Preferences for Self-Improving Reward Models
Alizée Pace, Jonathan Mallinson, Eric Malmi +2
The success of reinforcement learning from human feedback (RLHF) in language model alignment is strongly dependent on the quality of the underlying reward model. In this paper, we…
Effects of Research Paper Promotion via ArXiv and X
Chhandak Bagchi, Eric Malmi, Przemyslaw Grabowicz
In the evolving landscape of scientific publishing, it is important to understand the drivers of high-impact research, to equip scientists with actionable strategies to enhance the…