3 papers
cs.IR2026
Objective Shaping with Hard Negatives: Windowed Partial AUC Optimization for RL-based LLM Recommenders
Wentao Shi, Qifan Wang, Chen Chen +7
Reinforcement learning (RL) effectively optimizes Large Language Model (LLM)-based recommenders by contrasting positive and negative items. Empirically, training with beam-search n…
cs.LG2024
Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression
Yixiu Mao, Qi Wang, Chen Chen +2
In offline reinforcement learning (RL), addressing the out-of-distribution (OOD) action issue has been a focus, but we argue that there exists an OOD state issue that also impairs…
cs.AI2024
Hokoff: Real Game Dataset from Honor of Kings and its Offline Reinforcement Learning Benchmarks
Yun Qu, Boyuan Wang, Jianzhun Shao +15
The advancement of Offline Reinforcement Learning (RL) and Offline Multi-Agent Reinforcement Learning (MARL) critically depends on the availability of high-quality, pre-collected o…