3 papers
cs.CL2026
GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning
Outongyi Lv, Yuanwei Zhang, Xiaoqun Zhang
Reinforcement learning (RL), particularly RL with Verifiable Rewards (RLVR), has recently emerged as a central paradigm for enhancing large language models' (LLMs) reasoning abilit…
cs.AI2026
Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index
Outongyi Lv, Yanzhao Zheng, Yuanwei Zhang +5
Reinforcement learning (RL) has become a powerful tool for propelling Large Language Models (LLMs) beyond imitation-based training towards more robust reasoning capabilities. Among…
cs.LG2025
Improved Offline Reinforcement Learning via Quantum Metric Encoding
Outongyi Lv, Yewei Yuan, Nana Liu
Reinforcement learning (RL) with limited samples is common in real-world applications. However, offline RL performance under this constraint is often suboptimal. We consider an alt…