16 citations · 23 across the 4 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2023
ULMA: Unified Language Model Alignment with Human Demonstration and Point-wise Preference
Tianchi Cai, Xierui Song, Jiyan Jiang +3
Aligning language models to human expectations, e.g., being helpful and harmless, has become a pressing challenge for large language models. A typical alignment procedure consists…
cs.LG2023★ 1 cited
Marketing Budget Allocation with Offline Constrained Deep Reinforcement Learning
Tianchi Cai, Jiyan Jiang, Wenpeng Zhang +7
We study the budget allocation problem in online marketing campaigns that utilize previously collected offline data. We first discuss the long-term effect of optimizing marketing b…
cs.LG2023
Model-free Reinforcement Learning with Stochastic Reward Stabilization for Recommender Systems
Tianchi Cai, Shenliao Bao, Jiyan Jiang +5
Model-free RL-based recommender systems have recently received increasing research attention due to their capability to handle partial feedback and long-term rewards. However, most…