1 paper · 1 filter
Adrian Hutter
We consider a scenario in which two reinforcement learning agents repeatedly play a matrix game against each other and update their parameters after each round. The agents' decisio…