most citedPre-training Generative Recommender with Multi-Identifier Item Tokenization

1 citations · 1 across the 6 of their papers we have counts for

collaborators

8 papers

cs.LG2026

Towards a Theoretical Understanding to the Generalization of RLHF

Zhaochun Li, Mingyang Yi, Yue Wang +2

Reinforcement Learning from Human Feedback (RLHF) and its variants have emerged as the dominant approaches for aligning Large Language Models with human intent. While empirically e…

cs.LG2025

Stabilizing Policy Gradient Methods via Reward Profiling

Shihab Ahmed, El Houcine Bergou, Aritra Dutta +1

Policy gradient methods, which have been extensively studied in the last decade, offer an effective and efficient framework for reinforcement learning problems. However, their perf…

cs.IR20251 cited

Pre-training Generative Recommender with Multi-Identifier Item Tokenization

Bowen Zheng, Enze Liu, Zhongfu Chen +4

Generative recommendation autoregressively generates item identifiers to recommend potential items. Existing methods typically adopt a one-to-one mapping strategy, where each item…

cs.LG2025

DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management

Xuerui Su, Liya Guo, Yue Wang +4

Inference scaling further accelerates Large Language Models (LLMs) toward Artificial General Intelligence (AGI), with large-scale Reinforcement Learning (RL) to unleash long Chain-…

cs.LG2025

Symbolic Representation for Any-to-Any Generative Tasks

Jiaqi Chen, Xiaoye Zhu, Yue Wang +9

We propose a symbolic generative task description language and a corresponding inference engine capable of representing arbitrary multimodal tasks as structured symbolic flows. Unl…

cs.LG2025

Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Xuerui Su, Shufang Xie, Guoqing Liu +7

Recently, Large Language Models (LLMs) have rapidly evolved, approaching Artificial General Intelligence (AGI) while benefiting from large-scale reinforcement learning to enhance H…