2 papers
cs.CL2025
MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment
Tianze Wang, Dongnan Gui, Yifan Hu +2
Reinforcement Learning from Human Feedback (RLHF) has shown promise in aligning large language models (LLMs). Yet its reliance on a singular reward model often overlooks the divers…
cs.CL2024
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
Fan Nie, Xiaotian Hou, Shuhang Lin +3
The propensity of Large Language Models (LLMs) to generate hallucinations and non-factual content undermines their reliability in high-stakes domains, where rigorous control over T…