3 papers
cs.LG2025
Probabilistic Uncertain Reward Model
Wangtao Sun, Xiang Cheng, Xing Yu +5
Reinforcement learning from human feedback (RLHF) is a critical technique for training large language models. However, conventional reward models based on the Bradley-Terry model (…
cs.CL2025
Exploiting Contextual Knowledge in LLMs through V-usable Information based Layer Enhancement
Xiaowei Yuan, Zhao Yang, Ziyang Huang +5
Large Language Models (LLMs) have demonstrated remarkable capabilities in various tasks, yet they often struggle with context-faithfulness generations that properly reflect context…
cs.CL2024
Improving Zero-shot LLM Re-Ranker with Risk Minimization
Xiaowei Yuan, Zhao Yang, Yequan Wang +2
In the Retrieval-Augmented Generation (RAG) system, advanced Large Language Models (LLMs) have emerged as effective Query Likelihood Models (QLMs) in an unsupervised way, which re-…