3 papers
cs.CL2026
Negative Advantages Is a Double-Edged Sword: Calibrating advantages in GRPO for Search Agents
Jiayi Wu, Ruobing Xie, Zeqian Huang +6
Search agents achieve strong question-answering performance through multi-turn interactions with search engines, with Group Relative Policy Optimization (GRPO) being a widely used…
cs.IR2023
Multi-Feature Integration for Perception-Dependent Examination-Bias Estimation
Xiaoshu Chen, Xiangsheng Li, Kunliang Wei +4
Eliminating examination bias accurately is pivotal to apply click-through data to train an unbiased ranking model. However, most examination-bias estimators are limited to the hypo…
cs.IR2023
Pretraining De-Biased Language Model with Large-scale Click Logs for Document Ranking
Xiangsheng Li, Xiaoshu Chen, Kunliang Wei +4
Pre-trained language models have achieved great success in various large-scale information retrieval tasks. However, most of pretraining tasks are based on counterfeit retrieval da…