Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AutoTool: Automatic Scaling of Tool-Use Capabilities in RL via Decoupled Entropy Constraints
Yirong Zeng, Xiao Ding, Yufei Liu +9
Tool use represents a critical capability for AI agents, with recent advances focusing on leveraging reinforcement learning (RL) to scale up the explicit reasoning process to achie…
cs.AI2025
Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
Zhangying Feng, Qianglong Chen, Ning Lu +6
The development of reasoning capabilities represents a critical frontier in large language models (LLMs) research, where reinforcement learning (RL) and process reward models (PRMs…