1 paper
Jelena Markovic-Voronov, Wenhui Zhu, Bo Long +5
We introduce a principled probabilistic framework for reward-guided decoding in large language models, addressing the limitations of standard decoding methods that optimize token-l…