2 papers
cs.LG2025
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Zhenghai Xue, Longtao Zheng, Qian Liu +4
Large Language Models (LLMs) can significantly improve their reasoning capabilities by interacting with external tools, a paradigm known as Tool-Integrated Reasoning (TIR). However…
cs.LG2025
Logit Dynamics in Softmax Policy Gradient Methods
Yingru Li
We analyzes the logit dynamics of softmax policy gradient methods. We derive the exact formula for the L2 norm of the logit update vector: $$ \|Î\mathbf{z}\|_2 \propto \sqrt{1-2P_…