1 citations · 1 across the 1 of their papers we have counts for
1 paper
Ruijian Zha, Bojun Liu
Recent advances in reinforcement learning, such as Dynamic Sampling Policy Optimization (DAPO), show strong performance when paired with large language models (LLMs). Motivated by…