activity
20242026
collaborators

7 papers

cs.CL2026

Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

Sinie van der Ben, Raphaël Baur, Yannick Metz +1

Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally influence behavior, and exhibit geometry mirr…

cs.LG2026

MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference

Raphaël Baur, Yannick Metz, Maria Gkoulta +3

Reward learning typically relies on a single feedback type or combines multiple feedback types using manually weighted loss terms. Currently, it remains unclear how to jointly lear…

cs.LG2026

Evaluating Autoencoders for Parametric and Invertible Multidimensional Projections

Frederik L. Dennig, Nina Geyer, Daniela Blumberg +2

Recently, neural networks have gained attention for creating parametric and invertible multidimensional data projections. Parametric projections allow for embedding previously unse…

cs.LG2025

ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning

Timo Kaufmann, Yannick Metz, Daniel Keim +1

Binary choices, as often used for reinforcement learning from human feedback (RLHF), convey only the direction of a preference. A person may choose apples over oranges and bananas…

cs.LG2025

Reward Learning from Multiple Feedback Types

Yannick Metz, András Geiszl, Raphaël Baur +1

Learning rewards from preference feedback has become an important tool in the alignment of agentic models. Preference-based feedback, often implemented as a binary comparison betwe…

cs.LG2025

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework

Yannick Metz, David Lindner, Raphaël Baur +1

Reinforcement Learning from Human feedback (RLHF) has become a powerful tool to fine-tune or train agentic machine learning models. Similar to how humans interact in social context…