3 papers
cs.LG2026
TrasMuon: Trust-Region Adaptive Scaling for Orthogonalized Momentum Optimizers
Peng Cheng, Jiucheng Zang, Qingnan Li +6
Muon-style optimizers leverage Newton-Schulz (NS) iterations to orthogonalize updates, yielding update geometries that often outperform Adam-series methods. However, this orthogona…
cs.AI2026
Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
Michael R. Metel, Yufei Cui, Boxing Chen +1
Sequential test-time scaling is a promising training-free method to improve large reasoning model accuracy, but as currently implemented, significant limitations have been observed…
cs.LG2025
GRPO-: Credit Assignment improves LLM Reasoning
Prasanna Parthasarathi, Mathieu Reymond, Boxing Chen +2
Large language models (LLMs) are increasingly deployed for tasks requiring complex reasoning, prompting significant interest in improving their reasoning abilities through post-tra…