3 citations · 3 across the 8 of their papers we have counts for
1 paper · 1 filter
Tao Wang, Suhang Zheng, Xiaoxiao Xu
Multi-step agentic reinforcement learning benefits from fine-grained credit assignment, yet existing approaches offer limited options: critic-free methods like GRPO assign a unifor…