collaborators

6 papers

cs.CL2026

Unveiling the Entropy Dynamics of Chain-of-Thought Reasoning

Ting Xu, Xu He, Yupu Lu +4

This paper investigates the entropy dynamics of Chain-of-Thought (CoT) and uncovers a consistent two-phase structure: an Uncertainty Region of exploration transitioning sharply to…

cs.LG2026

Ratio-Variance Regularized Policy Optimization

Yu Luo, Shuo Han, Yihan Hu +5

Standard on-policy reinforcement learning relies on heuristic clipping to enforce trust regions, but this mechanism imposes a severe cost by indiscriminately truncating high-return…

cs.AI2026

: Unlocking LLM Reasoning via Reinforcement Learning with Re-solving

Pinzheng Wang, Shuli Xu, Juntao Li +4

Reinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning performance of large language models (LLMs) by increasing test-time compute. Howe…

cs.LG2026

Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning

Yu Luo, Shuo Han, Yihan Hu +2

On-policy reinforcement learning (RL), particularly Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), has become the dominant paradigm for fine-tuni…

cs.LG2024

Reinforced In-Context Black-Box Optimization

Lei Song, Chenxiao Gao, Ke Xue +5

Black-Box Optimization (BBO) has found successful applications in many fields of science and engineering. Recently, there has been a growing interest in meta-learning particular co…

cs.CV2024

iVideoGPT: Interactive VideoGPTs are Scalable World Models

Jialong Wu, Shaofeng Yin, Ningya Feng +4

World models empower model-based agents to interactively explore, reason, and plan within imagined environments for real-world decision-making. However, the high demand for interac…