1 paper
Hyobin Park, Taeseop Kim, Dong-Geol Choi
Self-play reinforcement learning has shown strong performance in domains with formally verifiable structure, such as mathematics and coding, where both problem generation and rewar…