4 papers
Better Estimation of the Kullback--Leibler Divergence Between Language Models
Afra Amini, Tim Vieira, Ryan Cotterell
Estimating the Kullback--Leibler (KL) divergence between language models has many applications, e.g., reinforcement learning from human feedback (RLHF), interpretability, and knowl…
Syntactic Control of Language Models by Posterior Inference
Vicky Xefteri, Tim Vieira, Ryan Cotterell +1
Controlling the syntactic structure of text generated by language models is valuable for applications requiring clarity, stylistic consistency, or interpretability, yet it remains…
Variational Best-of-N Alignment
Afra Amini, Tim Vieira, Elliott Ash +1
Best-of-N (BoN) is a popular and effective algorithm for aligning language models to human preferences. The algorithm works as follows: at inference time, N samples are drawn from…
Reverse-Engineering the Reader
Samuel Kiegeland, Ethan Gotlieb Wilcox, Afra Amini +2
Numerous previous studies have sought to determine to what extent language models, pretrained on natural language text, can serve as useful models of human cognition. In this paper…