5 papers · 1 filter
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou +4
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot gene…
Improving Generative Ad Text on Facebook using Reinforcement Learning
Daniel R. Jiang, Alex Nikulkov, Yu-Chia Chen +2
Generative artificial intelligence (AI), in particular large language models (LLMs), is poised to drive transformative economic change. LLMs are pre-trained on vast text data to le…
Aligned Multi Objective Optimization
Yonathan Efroni, Ben Kretzu, Daniel Jiang +4
To date, the multi-objective optimization literature has mainly focused on conflicting objectives, studying the Pareto front, or requiring users to balance tradeoffs. Yet, in machi…
Exploiting Structure in Offline Multi-Agent RL: The Benefits of Low Interaction Rank
Wenhao Zhan, Scott Fujimoto, Zheqing Zhu +3
We study the problem of learning an approximate equilibrium in the offline multi-agent reinforcement learning (MARL) setting. We introduce a structural assumption -- the interactio…
Pearl: A Production-ready Reinforcement Learning Agent
Zheqing Zhu, Rodrigo de Salvo Braz, Jalaj Bhandari +12
Reinforcement learning (RL) is a versatile framework for optimizing long-term goals. Although many real-world problems can be formalized with RL, learning and deploying a performan…