4 papers
Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling
Alex Nikulkov
Reward models in RLHF are trained to score only the final token of a response - a choice that discards rich signal from every intermediate position and produces models whose token-…
Improving Generative Ad Text on Facebook using Reinforcement Learning
Daniel R. Jiang, Alex Nikulkov, Yu-Chia Chen +2
Generative artificial intelligence (AI), in particular large language models (LLMs), is poised to drive transformative economic change. LLMs are pre-trained on vast text data to le…
Pearl: A Production-ready Reinforcement Learning Agent
Zheqing Zhu, Rodrigo de Salvo Braz, Jalaj Bhandari +12
Reinforcement learning (RL) is a versatile framework for optimizing long-term goals. Although many real-world problems can be formalized with RL, learning and deploying a performan…
Offline Reinforcement Learning for Optimizing Production Bidding Policies
Dmytro Korenkevych, Frank Cheng, Artsiom Balakir +5
The online advertising market, with its thousands of auctions run per second, presents a daunting challenge for advertisers who wish to optimize their spend under a budget constrai…