45 citations · 231 across the 25 of their papers we have counts for
3 papers · 1 filter
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
Stephane Hatgis-Kessell, Emma Brunskill
We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.e., when can we replace classical RL algorith…
Trading off rewards and errors in multi-armed bandits
Akram Erraqabi, Alessandro Lazaric, Michal Valko +2
In multi-armed bandits, the most-explored arms are the most informative, while reward maximization typically pulls only the best arm. We study the tradeoff between identifying arm…
GIANTS: Generative Insight Anticipation from Scientific Literature
Joy He-Yueya, Anikait Singh, Ge Gao +5
Scientific breakthroughs often emerge from synthesizing prior ideas into novel contributions. While language models (LMs) show promise in scientific discovery, their ability to per…