3 papers
cs.CL2026
Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning
Weiyu Ma, Yongcheng Zeng, Yan Song +4
Reinforcement Learning (RL) has achieved impressive success in post-training Large Language Models (LLMs) and Vision-Language Models (VLMs), with on-policy algorithms such as PPO,…
cs.AI2025
A Survey on Large Language Models for Mathematical Reasoning
Peng-Yuan Wang, Tian-Shuo Liu, Chenyang Wang +8
Mathematical reasoning has long represented one of the most fundamental and challenging frontiers in artificial intelligence research. In recent years, large language models (LLMs)…
cs.LG2024
Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning
Xu-Hui Liu, Tian-Shuo Liu, Shengyi Jiang +4
Combining offline and online reinforcement learning (RL) techniques is indeed crucial for achieving efficient and safe learning where data acquisition is expensive. Existing method…