activity
20242026
collaborators

20 papers

cs.IR2026

Token-Level Credit Assignment Optimization for Generative Document Retrieval

Xinpeng Zhao, Yang Liu, Ran Chen +6

Generative retrieval models perform document retrieval by autoregressively generating document identifiers (DocIDs). This process naturally forms a sequential decision problem, whe…

cs.CL2026

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

Ziyou Hu, Zhengliang Shi, Minghang Zhu +5

Reward models (RMs) have become essential for aligning large language models (LLMs), serving as scalable proxies for human evaluation in both training and inference. However, exist…

cs.LG2026

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning

Zhao Yang, Yuxuan Jiang, Ting-Chih Chen +18

Reinforcement learning (RL) has become central to LLM post-training, yet the methods that dominate current pipelines, PPO and GRPO, represent only a narrow slice of what RL offers.…

cs.CL2026

Disentangling Knowledge Representations for Large Language Model Editing

Mengqi Zhang, Zisheng Zhou, Xiaotian Ye +4

Knowledge Editing has emerged as a promising solution for efficiently updating embedded knowledge in large language models (LLMs). While existing approaches demonstrate effectivene…

cs.IR2026

SA-CAISR: Stage-Adaptive and Conflict-Aware Incremental Sequential Recommendation

Xiaomeng Song, Xinru Wang, Hanbing Wang +4

Sequential recommendation (SR) aims to predict a user's next action by learning from their historical interaction sequences. In real-world applications, these models require period…

cs.IR2026

Curriculum Approximate Unlearning for Session-based Recommendation

Liu Yang, Zhaochun Ren, Ziqi Zhao +7

Approximate unlearning for session-based recommendation refers to eliminating the influence of specific training samples from the recommender without retraining of (sub-)models. Gr…