collaborators

8 papers

cs.IR2026

GateSID: Adaptive Gating for Balancing Semantic and Collaborative Signals in Recommendation

Hai Zhu, Yantao Yu, Lei Shen +2

In cold-start scenarios, the scarcity of collaborative signals for new items exacerbates the Matthew effect, undermining platform diversity and posing a persistent challenge in pra…

cs.LG2026

Uncertainty-Aware Reward Modeling for Stable RLHF

Licheng Pan, Haocheng Yang, Haoxuan Li +7

Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies to maximize predicted rewards. H…

cs.LG2026

Optimal Transport for LLM Reward Modeling from Noisy Preference

Licheng Pan, Haochen Yang, Haoxuan Li +8

Reward models are fundamental to Reinforcement Learning from Human Feedback (RLHF), yet real-world datasets are inevitably corrupted by noisy preference. Conventional training obje…

cs.IR2026

Beyond Dense Connectivity: Explicit Sparsity for Scalable Recommendation

Yantao Yu, Sen Qiao, Lei Shen +2

Recent progress in scaling large models has motivated recommender systems to increase model depth and capacity to better leverage massive behavioral data. However, recommendation i…

cs.IR2026

SIGMA: A Semantic-Grounded Instruction-Driven Generative Multi-Task Recommender at AliExpress

Yang Yu, Lei Kou, Huaikuan Yi +6

With the rapid evolution of Large Language Models (LLMs), generative recommendation is gradually reshaping the paradigm of recommender systems. However, most existing methods remai…

cs.CL2026

ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment

Hao Wang, Haocheng Yang, Licheng Pan +7

Reward modeling represents a long-standing challenge in reinforcement learning from human feedback (RLHF) for aligning language models. Current reward modeling is heavily contingen…