Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
LiPO: Listwise Preference Optimization through Learning-to-Rank
Tianqi Liu, Zhen Qin, Junru Wu +9
Aligning language models (LMs) with curated human feedback is critical to control their behaviors in real-world applications. Several recent policy optimization methods, such as DP…
cs.CL2023
Self-Evaluation Improves Selective Generation in Large Language Models
Jie Ren, Yao Zhao, Tu Vu +2
Safe deployment of large language models (LLMs) may benefit from a reliable method for assessing their generated content to determine when to abstain or to selectively generate. Wh…