4 papers
Owen-Shapley Policy Optimization: A Principled RL Algorithm for Generative Search LLMs
Abhijnan Nath, Alireza Bagheri Garakani, Tianchen Zhou +3
Large language models are increasingly trained via reinforcement learning for personalized recommendation tasks, but standard methods like GRPO rely on sparse, sequence-level rewar…
Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach
Fnu Hairi, Jiao Yang, Tianchen Zhou +6
In many multi-objective reinforcement learning (MORL) applications, being able to systematically explore the Pareto-stationary solutions under multiple non-convex reward objectives…
Large Language Model-Enhanced Reinforcement Learning for Diverse and Novel Recommendations
Jiin Woo, Alireza Bagheri Garakani, Tianchen Zhou +2
In recommendation systems, diversity and novelty are essential for capturing varied user preferences and encouraging exploration, yet many systems prioritize click relevance. While…
PCL: Prompt-based Continual Learning for User Modeling in Recommender Systems
Mingdai Yang, Fan Yang, Yanhui Guo +6
User modeling in large e-commerce platforms aims to optimize user experiences by incorporating various customer activities. Traditional models targeting a single task often focus o…