Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Reward Learning from Multiple Feedback Types
Yannick Metz, András Geiszl, Raphaël Baur +1
Learning rewards from preference feedback has become an important tool in the alignment of agentic models. Preference-based feedback, often implemented as a binary comparison betwe…
cs.LG2024
Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework
Yannick Metz, David Lindner, Raphaël Baur +1
Reinforcement Learning from Human feedback (RLHF) has become a powerful tool to fine-tune or train agentic machine learning models. Similar to how humans interact in social context…