4 papers
Open Problems in Constitutional Preference Reconstruction
Eleanor Clifford, Michael Amir, Arduin Findeis +2
Pairwise preference data is widely used for training and evaluating language models (e.g., RLHF), but each datapoint records a \emph{choice}, not the rationale behind it. Methods s…
Feedback Forensics: A Toolkit to Measure AI Personality
Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier +1
Some traits making a "good" AI model are hard to describe upfront. For example, should responses be more polite or more casual? Such traits are sometimes summarized as model charac…
Learning from Preferences and Mixed Demonstrations in General Settings
Jason R Brown, Carl Henrik Ek, Robert D Mullins
Reinforcement learning is a general method for learning in sequential settings, but it can often be difficult to specify a good reward function when the task is complex. In these c…
Inverse Constitutional AI: Compressing Preferences into Principles
Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier +2
Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options,…