3 papers
cs.CV2026
Arena as Offline Reward: Efficient Fine-Grained Preference Optimization for Diffusion Models
Zhikai Li, Yue Zhao, Edward Zhongwei Zhang +4
Reinforcement learning from human feedback (RLHF) effectively promotes preference alignment of text-to-image (T2I) diffusion models. To improve computational efficiency, direct pre…
cs.IR2026
Effective Knowledge Transfer for Multi-Task Recommendation Models
Guohao Cai, Jun Yuan, Zhenhua Dong
The conversion rate (CVR) is a crucial metric for evaluating the effectiveness of platforms, as it quantifies the alignment of content with audience preferences. However, the limit…
cs.IR2024
RecSys Arena: Pair-wise Recommender System Evaluation with Large Language Models
Zhuo Wu, Qinglin Jia, Chuhan Wu +4
Evaluating the quality of recommender systems is critical for algorithm design and optimization. Most evaluation methods are computed based on offline metrics for quick algorithm e…