1 paper
Agnibh Karmakar, Mayur Parvatikar, Shreyash Dhoot +5
Decode-time alignment methods steer a frozen language model by scoring candidate continuations with an external reward and selecting the maximiser. We argue that this shared design…