4 papers · 1 filter
Extreme Region Policy Distillation
Changyu Chen, Xiting Wang, Rui Yan
Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories…
Controlled LLM Training on Spectral Sphere
Tian Xie, Haoming Luo, Haoyu Tang +9
Scaling large models requires optimization strategies that ensure rapid convergence grounded in stability. Maximal Update Parametrization (P) provides a theoretical s…
Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
Mingkuan Zhao, Wentao Hu, Jiayin Wang +5
The design of Large Language Models (LLMs) has long been hampered by a fundamental conflict within their core attention mechanism: its remarkable expressivity is built upon a compu…
Semi-Offline Reinforcement Learning for Optimized Text Generation
Changyu Chen, Xiting Wang, Yiqiao Jin +5
In reinforcement learning (RL), there are two major settings for interacting with the environment: online and offline. Online methods explore the environment at significant time co…