collaborators

7 papers

cs.LG2026

Extreme Region Policy Distillation

Changyu Chen, Xiting Wang, Rui Yan

Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories…

cs.LG2026

Semi-Offline Reinforcement Learning for Optimized Text Generation

Changyu Chen, Xiting Wang, Yiqiao Jin +5

In reinforcement learning (RL), there are two major settings for interacting with the environment: online and offline. Online methods explore the environment at significant time co…

cs.LG2026

Controlled LLM Training on Spectral Sphere

Tian Xie, Haoming Luo, Haoyu Tang +9

Scaling large models requires optimization strategies that ensure rapid convergence grounded in stability. Maximal Update Parametrization (P) provides a theoretical…

cs.LG2025

Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off

Mingkuan Zhao, Wentao Hu, Jiayin Wang +5

The design of Large Language Models (LLMs) has long been hampered by a fundamental conflict within their core attention mechanism: its remarkable expressivity is built upon a compu…

cs.AI2025

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards

Jixiang Hong, Yiran Zhang, Guanzhong Wang +3

Building upon large language models (LLMs), recent large multimodal models (LMMs) unify cross-model understanding and generation into a single framework. However, LMMs still strugg…

cs.AI2025

RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing

Jianxing Liao, Tian Zhang, Xiao Feng +6

Large language models are extensively utilized in creative writing applications. Creative writing requires a balance between subjective writing quality (e.g., literariness and emot…