3 papers
cs.LG2026
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Zhao Yang, Yuxuan Jiang, Ting-Chih Chen +18
Reinforcement learning (RL) has become central to LLM post-training, yet the methods that dominate current pipelines, PPO and GRPO, represent only a narrow slice of what RL offers.…
cs.LG2024
Rocket Landing Control with Random Annealing Jump Start Reinforcement Learning
Yuxuan Jiang, Yujie Yang, Zhiqian Lan +6
Rocket recycling is a crucial pursuit in aerospace technology, aimed at reducing costs and environmental impact in space exploration. The primary focus centers on rocket landing co…
cs.LG2024
Diffusion Actor-Critic with Entropy Regulator
Yinuo Wang, Likun Wang, Yuxuan Jiang +8
Reinforcement learning (RL) has proven highly effective in addressing complex decision-making and control tasks. However, in most traditional RL algorithms, the policy is typically…