Showing 2026Show all
3 papers · 1 filter
cs.AI2026
Targeted Exploration via Unified Entropy Control for Reinforcement Learning
Chen Wang, Lai Wei, Yanzhi Zhang +5
Recent advances in reinforcement learning (RL) have improved the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs). However, the widely used…
cs.DC2026
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
Xu Bai, Muhammed Tawfiqul Islam, Chen Wang +1
Pipeline parallelism (PP) is widely used to partition layers of large language models (LLMs) across GPUs, enabling scalable inference for large models. However, existing systems re…
cs.LG2026
Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off
Zhaochun Li, Chen Wang, Jionghao Bai +4
The exploration-exploitation (EE) trade-off is a central challenge in reinforcement learning (RL) for large language models (LLMs). With Group Relative Policy Optimization (GRPO),…