most citedPre-training Generative Recommender with Multi-Identifier Item Tokenization

1 citations · 1 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

Towards a Theoretical Understanding to the Generalization of RLHF

Zhaochun Li, Mingyang Yi, Yue Wang +2

Reinforcement Learning from Human Feedback (RLHF) and its variants have emerged as the dominant approaches for aligning Large Language Models with human intent. While empirically e…

cs.LG2026

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment

Xiuyu Li, Jinkai Zhang, Mingyang Yi +4

Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complicated training process. To addres…

cs.LG2025

Stabilizing Policy Gradient Methods via Reward Profiling

Shihab Ahmed, El Houcine Bergou, Aritra Dutta +1

Policy gradient methods, which have been extensively studied in the last decade, offer an effective and efficient framework for reinforcement learning problems. However, their perf…

cs.LG2025

DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management

Xuerui Su, Liya Guo, Yue Wang +4

Inference scaling further accelerates Large Language Models (LLMs) toward Artificial General Intelligence (AGI), with large-scale Reinforcement Learning (RL) to unleash long Chain-…

cs.LG2025

Symbolic Representation for Any-to-Any Generative Tasks

Jiaqi Chen, Xiaoye Zhu, Yue Wang +9

We propose a symbolic generative task description language and a corresponding inference engine capable of representing arbitrary multimodal tasks as structured symbolic flows. Unl…

cs.LG2025

Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Xuerui Su, Shufang Xie, Guoqing Liu +7

Recently, Large Language Models (LLMs) have rapidly evolved, approaching Artificial General Intelligence (AGI) while benefiting from large-scale reinforcement learning to enhance H…