4 citations · 6 across the 12 of their papers we have counts for
1 paper · 2 filters
Tianyi Hu, Qingxu Fu, Yanxi Chen +2
Reinforcement learning (RL) has emerged as the predominant paradigm for training large language model (LLM)-based AI agents. However, existing backbone RL algorithms lack verified…