4 papers · 1 filter
Adaptive Decoding via Latent Preference Optimization
Shehzaad Dhuliawala, Ilia Kulikov, Ping Yu +4
During language model decoding, it is known that using higher temperature sampling gives more creative responses, while lower temperatures are more factually accurate. However, suc…
Self-Taught Evaluators
Tianlu Wang, Ilia Kulikov, Olga Golovneva +7
Model-based evaluation is at the heart of successful model development -- as a reward model for training, and as a replacement for human evaluation. To train such evaluators, the s…
Distilling System 2 into System 1
Ping Yu, Jing Xu, Jason Weston +1
Large language models (LLMs) can spend extra compute during inference to generate intermediate thoughts, which helps to produce better final responses. Since Chain-of-Thought (Wei…
Following Length Constraints in Instructions
Weizhe Yuan, Ilia Kulikov, Ping Yu +4
Aligned instruction following models can better fulfill user requests than their unaligned counterparts. However, it has been shown that there is a length bias in evaluation of suc…