2 citations · 2 across the 8 of their papers we have counts for
Showing 2026 · cs.LGShow all
2 papers · 2 filters
cs.LG2026
Extreme Region Policy Distillation
Changyu Chen, Xiting Wang, Rui Yan
Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories…
cs.LG2026
Controlled LLM Training on Spectral Sphere
Tian Xie, Haoming Luo, Haoyu Tang +9
Scaling large models requires optimization strategies that ensure rapid convergence grounded in stability. Maximal Update Parametrization (P) provides a theoretical s…