3 papers
cs.LG2026
On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation
Anton Baumann, Akmal Ashirmatov, Leo Schmidt-Traub +4
On-policy self-distillation provides dense, token-level supervision by conditioning a model on privileged information and distilling the resulting teacher distribution back into th…
cs.CL2026
Aligning Language Models from User Interactions
Thomas Kleine Buening, Jonas Hübotter, Barna Pásztor +3
Multi-turn user interactions are among the most abundant data produced by language models, yet we lack effective methods to learn from them. While typically discarded, these intera…
cs.LG2026
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
Raphaël Baur, Yannick Metz, Maria Gkoulta +3
Reward learning typically relies on a single feedback type or combines multiple feedback types using manually weighted loss terms. Currently, it remains unclear how to jointly lear…