works on

From the 1 of 8 linked papers with an AI index.

activity
20242026
collaborators

8 papers

cs.CL2026

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

Patrik Wolf, Thomas Kleine Buening, Andreas Krause +1

The paper investigates whether large language models give probability estimates that satisfy the law of total probability when prompted for subpopulations, revealing frequent viola…

cs.LG2026

Reinforcement Learning via Self-Distillation

Jonas Hübotter, Frederike Lübeck, Lejs Behric +8

Large language models are increasingly post-trained with reinforcement learning in verifiable domains such as code and math. Yet, current methods for reinforcement learning with ve…

cs.LG2026

Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding

Daqian Shao, Thomas Kleine Buening, Marta Kwiatkowska

We propose a general framework for causal Imitation Learning (IL) with hidden confounders, which subsumes several existing settings. Our framework accounts for two types of hidden…

cs.LG2025

Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game

Barna Pásztor, Thomas Kleine Buening, Andreas Krause

We introduce Stackelberg Learning from Human Feedback (SLHF), a new framework for preference optimization. SLHF frames the alignment problem as a sequential-move game between two p…

cs.LG2025

Strategyproof Reinforcement Learning from Human Feedback

Thomas Kleine Buening, Jiarui Gan, Debmalya Mandal +1

We study Reinforcement Learning from Human Feedback (RLHF) in settings where multiple labelers may strategically misreport feedback to steer the learned policy toward their own pre…

cs.AI2025

A Minimax Approach to Ad Hoc Teamwork

Victor Villin, Thomas Kleine Buening, Christos Dimitrakakis

We propose a minimax-Bayes approach to Ad Hoc Teamwork (AHT) that optimizes policies against an adversarial prior over partners, explicitly accounting for uncertainty about partner…