activity
20242026
collaborators

6 papers

cs.LG2026

Calibrated Preference Learning: The Case of Label Ranking

Santo M. A. R. Thies, Viktor Bengs, Timo Kaufmann +2

Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively studied for classification and reg…

cs.LG2025

ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning

Timo Kaufmann, Yannick Metz, Daniel Keim +1

Binary choices, as often used for reinforcement learning from human feedback (RLHF), convey only the direction of a preference. A person may choose apples over oranges and bananas…

cs.CL2025

Feedback Forensics: A Toolkit to Measure AI Personality

Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier +1

Some traits making a "good" AI model are hard to describe upfront. For example, should responses be more polite or more casual? Such traits are sometimes summarized as model charac…

cs.AI2024

Problem Solving Through Human-AI Preference-Based Cooperation

Subhabrata Dutta, Timo Kaufmann, Goran Glavaš +7

While there is a widespread belief that artificial general intelligence (AGI) -- or even superhuman AI -- is imminent, complex problems in expert domains are far from being solved.…

cs.LG2024

OCALM: Object-Centric Assessment with Language Models

Timo Kaufmann, Jannis Blüml, Antonia Wüst +3

Properly defining a reward signal to efficiently train a reinforcement learning (RL) agent is a challenging task. Designing balanced objective functions from which a desired behavi…

cs.CL2024

Inverse Constitutional AI: Compressing Preferences into Principles

Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier +2

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options,…