21 citations · 21 across the 2 of their papers we have counts for
3 papers
Universal Length Generalization with Turing Programs
Kaiying Hou, David Brandfonbrener, Sham Kakade +2
Length generalization refers to the ability to extrapolate from short training sequences to long test sequences and is a challenge for current large language models. While prior wo…
Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning
Denis Yarats, David Brandfonbrener, Hao Liu +4
Recent progress in deep learning has relied on access to large and diverse datasets. Such data-driven progress has been less evident in offline reinforcement learning (RL), because…
Quantile Filtered Imitation Learning
David Brandfonbrener, William F. Whitney, Rajesh Ranganath +1
We introduce quantile filtered imitation learning (QFIL), a novel policy improvement operator designed for offline reinforcement learning. QFIL performs policy improvement by runni…