From the 1 of 8 linked papers with an AI index.
8 papers
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models
Patrik Wolf, Thomas Kleine Buening, Andreas Krause +1
The paper investigates whether large language models give probability estimates that satisfy the law of total probability when prompted for subpopulations, revealing frequent viola…
Reinforcement Learning via Self-Distillation
Jonas Hübotter, Frederike Lübeck, Lejs Behric +8
Large language models are increasingly post-trained with reinforcement learning in verifiable domains such as code and math. Yet, current methods for reinforcement learning with ve…
Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding
Daqian Shao, Thomas Kleine Buening, Marta Kwiatkowska
We propose a general framework for causal Imitation Learning (IL) with hidden confounders, which subsumes several existing settings. Our framework accounts for two types of hidden…
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
Barna Pásztor, Thomas Kleine Buening, Andreas Krause
We introduce Stackelberg Learning from Human Feedback (SLHF), a new framework for preference optimization. SLHF frames the alignment problem as a sequential-move game between two p…
Strategyproof Reinforcement Learning from Human Feedback
Thomas Kleine Buening, Jiarui Gan, Debmalya Mandal +1
We study Reinforcement Learning from Human Feedback (RLHF) in settings where multiple labelers may strategically misreport feedback to steer the learned policy toward their own pre…
A Minimax Approach to Ad Hoc Teamwork
Victor Villin, Thomas Kleine Buening, Christos Dimitrakakis
We propose a minimax-Bayes approach to Ad Hoc Teamwork (AHT) that optimizes policies against an adversarial prior over partners, explicitly accounting for uncertainty about partner…