1 paper
Thomas Pierrot, Guillaume Ligner, Scott Reed +6
We propose a novel reinforcement learning algorithm, AlphaNPI, that incorporates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero. NPI contributes structural bia…