2 citations · 2 across the 1 of their papers we have counts for
1 paper
Joseph Palermo, Johnny Ye, Alok Singh
We convert the DeepMind Mathematics Dataset into a reinforcement learning environment by interpreting it as a program synthesis problem. Each action taken in the environment adds a…