1 citations · 2 across the 8 of their papers we have counts for
1 paper · 1 filter
Hongliang Lu, Yuhang Wen, Pengyu Cheng +7
Reinforcement learning with verifiable rewards (RLVR) has become the mainstream technique for training LLM agents. However, RLVR highly depends on well-crafted task queries and cor…