2 papers
cs.AI2025
Establishing Reliability Metrics for Reward Models in Large Language Models
Yizhou Chen, Yawen Liu, Xuesi Wang +5
The reward model (RM) that represents human preferences plays a crucial role in optimizing the outputs of large language models (LLMs), e.g., through reinforcement learning from hu…
cs.IR2024
Residual Multi-Task Learner for Applied Ranking
Cong Fu, Kun Wang, Jiahua Wu +5
Modern e-commerce platforms rely heavily on modeling diverse user feedback to provide personalized services. Consequently, multi-task learning has become an integral part of their…