2 papers
cs.LG2026
Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling
Alex Nikulkov
Reward models in RLHF are trained to score only the final token of a response - a choice that discards rich signal from every intermediate position and produces models whose token-…
cs.LG2025
Improving Generative Ad Text on Facebook using Reinforcement Learning
Daniel R. Jiang, Alex Nikulkov, Yu-Chia Chen +2
Generative artificial intelligence (AI), in particular large language models (LLMs), is poised to drive transformative economic change. LLMs are pre-trained on vast text data to le…