6 papers
Unveiling the Entropy Dynamics of Chain-of-Thought Reasoning
Ting Xu, Xu He, Yupu Lu +4
This paper investigates the entropy dynamics of Chain-of-Thought (CoT) and uncovers a consistent two-phase structure: an Uncertainty Region of exploration transitioning sharply to…
Ratio-Variance Regularized Policy Optimization
Yu Luo, Shuo Han, Yihan Hu +5
Standard on-policy reinforcement learning relies on heuristic clipping to enforce trust regions, but this mechanism imposes a severe cost by indiscriminately truncating high-return…
: Unlocking LLM Reasoning via Reinforcement Learning with Re-solving
Pinzheng Wang, Shuli Xu, Juntao Li +4
Reinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning performance of large language models (LLMs) by increasing test-time compute. Howe…
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
Yu Luo, Shuo Han, Yihan Hu +2
On-policy reinforcement learning (RL), particularly Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), has become the dominant paradigm for fine-tuni…
Reinforced In-Context Black-Box Optimization
Lei Song, Chenxiao Gao, Ke Xue +5
Black-Box Optimization (BBO) has found successful applications in many fields of science and engineering. Recently, there has been a growing interest in meta-learning particular co…
iVideoGPT: Interactive VideoGPTs are Scalable World Models
Jialong Wu, Shaofeng Yin, Ningya Feng +4
World models empower model-based agents to interactively explore, reason, and plan within imagined environments for real-world decision-making. However, the high demand for interac…