Reinforcement Learning for Generative AI: State of the Art, Opportunities and Open Research Challenges
arXiv:2308.00031 · doi:10.1613/jair.1.15278
Abstract
Generative Artificial Intelligence (AI) is one of the most exciting developments in Computer Science of the last decade. At the same time, Reinforcement Learning (RL) has emerged as a very successful paradigm for a variety of machine learning tasks. In this survey, we discuss the state of the art, opportunities and open research questions in applying RL to generative AI. In particular, we will discuss three types of applications, namely, RL as an alternative way for generation without specified objectives; as a way for generating outputs while concurrently maximizing an objective function; and, finally, as a way of embedding desired characteristics, which cannot be easily captured by means of an objective function, into the generative process. We conclude the survey with an in-depth discussion of the opportunities and challenges in this fascinating emerging area.
Published in JAIR at https://www.jair.org/index.php/jair/article/view/15278
References in corpus (47)
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Training language models to follow instructions with human feedback
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- On the Opportunities and Risks of Foundation Models
- Zero-Shot Text-to-Image Generation
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- Gemini: A Family of Highly Capable Multimodal Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- LaMDA: Language Models for Dialog Applications
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Long Text Generation via Adversarial Training with Leaked Information
- Why We Need New Evaluation Metrics for NLG
- Improving alignment of dialogue agents via targeted human judgements
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation
- On the Creativity of Large Language Models
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
- Copyright in Generative Deep Learning
- Collective Constitutional AI: Aligning a Language Model with Public Input
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
- Fundamental Limitations of Alignment in Large Language Models
- Quark: Controllable Text Generation with Reinforced Unlearning
- Optimizing Prompts for Text-to-Image Generation
- Graph Constrained Reinforcement Learning for Natural Language Action Spaces
- Aligning Text-to-Image Models using Human Feedback
- Design Guidelines for Prompt Engineering Text-to-Image Generative Models
- End-to-End Offline Goal-Oriented Dialog Policy Learning via Policy Gradient
- Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models
- Self-collaboration Code Generation via ChatGPT
- Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
- Training Diffusion Models with Reinforcement Learning
- Preference Ranking Optimization for Human Alignment
- Offline RL for Natural Language Generation with Implicit Language Q Learning
- Execution-based Code Generation using Deep Reinforcement Learning
- Batch Policy Gradient Methods for Improving Neural Conversation Models
- ColdGANs: Taming Language GANs with Cautious Sampling Strategies
- Generative Cooperative Networks for Natural Language Generation
- Causal Confusion and Reward Misidentification in Preference-Based Reward Learning
- Neural Collage Transfer: Artistic Reconstruction via Material Manipulation
- Rainier: Reinforced Knowledge Introspector for Commonsense Question Answering