26 citations · 27 across the 5 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control
Xincheng Yao, Ruoqi Li, Cheng Chen +4
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a pivotal technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, the de f…
cs.LG2024★ 1 cited
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Huaijie Wang, Shibo Hao, Hanze Dong +4
Improving the multi-step reasoning ability of large language models (LLMs) with offline reinforcement learning (RL) is essential for quickly adapting them to complex tasks. While D…