4 papers
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Sinie van der Ben, Raphaël Baur, Yannick Metz +1
Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behavior, and exhibit geometry mirr…
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
Raphaël Baur, Yannick Metz, Maria Gkoulta +3
Reward learning typically relies on a single feedback type or combines multiple feedback types using manually weighted loss terms. Currently, it remains unclear how to jointly lear…
Reward Learning from Multiple Feedback Types
Yannick Metz, András Geiszl, Raphaël Baur +1
Learning rewards from preference feedback has become an important tool in the alignment of agentic models. Preference-based feedback, often implemented as a binary comparison betwe…
Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework
Yannick Metz, David Lindner, Raphaël Baur +1
Reinforcement Learning from Human feedback (RLHF) has become a powerful tool to fine-tune or train agentic machine learning models. Similar to how humans interact in social context…