2 papers
cs.LG2026
IFNSO: Iteration-Free Newton-Schulz Orthogonalization
Chen Hu, Qianxi Zhao, Xiaochen Yuan +4
The Newton-Schulz (NS) iteration has become a key technique for orthogonalization in optimizers such as Muon and for optimization on the Stiefel manifold. Despite its effectiveness…
cs.CL2025
ToolExpander: Extending the Frontiers of Tool-Using Reinforcement Learning to Weak LLMs
Fu Chen, Peng Wang, Xiyin Li +3
Training Large Language Models (LLMs) with Group Relative Policy Optimization (GRPO) encounters a significant challenge: models often fail to produce accurate responses, particular…