7 papers
Extreme Region Policy Distillation
Changyu Chen, Xiting Wang, Rui Yan
Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories…
Semi-Offline Reinforcement Learning for Optimized Text Generation
Changyu Chen, Xiting Wang, Yiqiao Jin +5
In reinforcement learning (RL), there are two major settings for interacting with the environment: online and offline. Online methods explore the environment at significant time co…
Controlled LLM Training on Spectral Sphere
Tian Xie, Haoming Luo, Haoyu Tang +9
Scaling large models requires optimization strategies that ensure rapid convergence grounded in stability. Maximal Update Parametrization (P) provides a theoretical…
Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
Mingkuan Zhao, Wentao Hu, Jiayin Wang +5
The design of Large Language Models (LLMs) has long been hampered by a fundamental conflict within their core attention mechanism: its remarkable expressivity is built upon a compu…
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
Jixiang Hong, Yiran Zhang, Guanzhong Wang +3
Building upon large language models (LLMs), recent large multimodal models (LMMs) unify cross-model understanding and generation into a single framework. However, LMMs still strugg…
RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing
Jianxing Liao, Tian Zhang, Xiao Feng +6
Large language models are extensively utilized in creative writing applications. Creative writing requires a balance between subjective writing quality (e.g., literariness and emot…