Reinforcement Learning from Human Feedback: Whose Culture, Whose Values, Whose Perspectives?
arXiv:2407.17482 · doi:10.1007/s13347-025-00861-0
Abstract
We argue for the epistemic and ethical advantages of pluralism in Reinforcement Learning from Human Feedback (RLHF) in the context of Large Language Models (LLM). Drawing on social epistemology and pluralist philosophy of science, we suggest ways in which RHLF can be made more responsive to human needs and how we can address challenges along the way. The paper concludes with an agenda for change, i.e. concrete, actionable steps to improve LLM development.
References in corpus (15)
- Overcoming catastrophic forgetting in neural networks
- Training language models to follow instructions with human feedback
- Deep reinforcement learning from human preferences
- Fine-Tuning Language Models from Human Preferences
- Constitutional AI: Harmlessness from AI Feedback
- Managing extreme AI risks amid rapid progress
- Fine-tuning language models to find agreement among humans with diverse preferences
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
- Discovering Language Model Behaviors with Model-Written Evaluations
- Targeting Solutions in Bayesian Multi-Objective Optimization: Sequential and Batch Versions
- Multi-step Jailbreaking Privacy Attacks on ChatGPT
- Towards a Benchmark for Scientific Understanding in Humans and Machines
- Epistemic Injustice in Generative AI
- On the Impossible Safety of Large AI Models
- Reinforcement Learning with Feedback from Multiple Humans with Diverse Skills