8 citations · 9 across the 3 of their papers we have counts for
1 paper · 1 filter
Ruize Zhang, Zelai Xu, Chengdong Ma +8
Self-play, a learning paradigm where agents iteratively refine their policies by interacting with historical or concurrent versions of themselves or other evolving agents, has show…