7 papers
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Sinie van der Ben, Raphaël Baur, Yannick Metz +1
Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behavior, and exhibit geometry mirr…
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
Raphaël Baur, Yannick Metz, Maria Gkoulta +3
Reward learning typically relies on a single feedback type or combines multiple feedback types using manually weighted loss terms. Currently, it remains unclear how to jointly lear…
Evaluating Autoencoders for Parametric and Invertible Multidimensional Projections
Frederik L. Dennig, Nina Geyer, Daniela Blumberg +2
Recently, neural networks have gained attention for creating parametric and invertible multidimensional data projections. Parametric projections allow for embedding previously unse…
ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning
Timo Kaufmann, Yannick Metz, Daniel Keim +1
Binary choices, as often used for reinforcement learning from human feedback (RLHF), convey only the direction of a preference. A person may choose apples over oranges and bananas…
Reward Learning from Multiple Feedback Types
Yannick Metz, András Geiszl, Raphaël Baur +1
Learning rewards from preference feedback has become an important tool in the alignment of agentic models. Preference-based feedback, often implemented as a binary comparison betwe…
Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework
Yannick Metz, David Lindner, Raphaël Baur +1
Reinforcement Learning from Human feedback (RLHF) has become a powerful tool to fine-tune or train agentic machine learning models. Similar to how humans interact in social context…