2 papers
cs.LG2024
Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation
Xiaoying Zhang, Jean-Francois Ton, Wei Shen +2
We introduce Adversarial Policy Optimization (AdvPO), a novel solution to the pervasive issue of reward over-optimization in Reinforcement Learning from Human Feedback (RLHF) for L…
cs.IR2024
Full Stage Learning to Rank: A Unified Framework for Multi-Stage Systems
Kai Zheng, Haijun Zhao, Rui Huang +6
The Probability Ranking Principle (PRP) has been considered as the foundational standard in the design of information retrieval (IR) systems. The principle requires an IR module's…