8 papers · 1 filter
Better Estimation of the Kullback--Leibler Divergence Between Language Models
Afra Amini, Tim Vieira, Ryan Cotterell
Estimating the Kullback--Leibler (KL) divergence between language models has many applications, e.g., reinforcement learning from human feedback (RLHF), interpretability, and knowl…
Syntactic Control of Language Models by Posterior Inference
Vicky Xefteri, Tim Vieira, Ryan Cotterell +1
Controlling the syntactic structure of text generated by language models is valuable for applications requiring clarity, stylistic consistency, or interpretability, yet it remains…
Variational Best-of-N Alignment
Afra Amini, Tim Vieira, Elliott Ash +1
Best-of-N (BoN) is a popular and effective algorithm for aligning language models to human preferences. The algorithm works as follows: at inference time, N samples are drawn from…
Reverse-Engineering the Reader
Samuel Kiegeland, Ethan Gotlieb Wilcox, Afra Amini +2
Numerous previous studies have sought to determine to what extent language models, pretrained on natural language text, can serve as useful models of human cognition. In this paper…
Direct Preference Optimization with an Offset
Afra Amini, Tim Vieira, Ryan Cotterell
Direct preference optimization (DPO) is a successful fine-tuning strategy for aligning large language models with human preferences without the need to train a reward model or empl…
Structured Voronoi Sampling
Afra Amini, Li Du, Ryan Cotterell
Gradient-based sampling algorithms have demonstrated their effectiveness in text generation, especially in the context of controlled text generation. However, there exists a lack o…