5 papers
Open Problems in Constitutional Preference Reconstruction
Eleanor Clifford, Michael Amir, Arduin Findeis +2
Pairwise preference data is widely used for training and evaluating language models (e.g., RLHF), but each datapoint records a \emph{choice}, not the rationale behind it. Methods s…
LLMs Contain Multitudes: How Deployment Context Reshapes Model-Level Preferences and Values
Filip Trhlik, Aoife O'Flynn, Angela Yu +2
Large language models (LLMs) are increasingly characterised in recent evaluation work as having stable, model-level preference and value systems. However, accompanying robustness c…
Feedback Forensics: A Toolkit to Measure AI Personality
Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier +1
Some traits making a "good" AI model are hard to describe upfront. For example, should responses be more polite or more casual? Such traits are sometimes summarized as model charac…
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
Arduin Findeis, Floris Weers, Guoli Yin +3
Pairwise preferences over model responses are widely collected to evaluate and provide feedback to large language models (LLMs). Given two alternative model responses to the same i…
Inverse Constitutional AI: Compressing Preferences into Principles
Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier +2
Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options,…