2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.LG2021★ 2 cited
A Reinforcement Learning Environment for Mathematical Reasoning via Program Synthesis
Joseph Palermo, Johnny Ye, Alok Singh
We convert the DeepMind Mathematics Dataset into a reinforcement learning environment by interpreting it as a program synthesis problem. Each action taken in the environment adds a…
cs.LG2019★ 1 cited
Detecting Spiky Corruption in Markov Decision Processes
Jason Mancuso, Tomasz Kisielewski, David Lindner +1
Current reinforcement learning methods fail if the reward function is imperfect, i.e. if the agent observes reward different from what it actually receives. We study this problem w…