5 papers
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
Ziqi Zhao, Xinyu Ma, Liu Yang +6
On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-policy rollouts. However, ex…
Reinforced Efficient Reasoning via Semantically Diverse Exploration
Ziqi Zhao, Zhaochun Ren, Jiahong Zou +9
Reinforcement learning with verifiable rewards (RLVR) has proven effective in enhancing the reasoning of large language models (LLMs). Monte Carlo Tree Search (MCTS)-based extensio…
Curriculum Approximate Unlearning for Session-based Recommendation
Liu Yang, Zhaochun Ren, Ziqi Zhao +7
Approximate unlearning for session-based recommendation refers to eliminating the influence of specific training samples from the recommender without retraining of (sub-)models. Gr…
Offline Trajectory Optimization for Offline Reinforcement Learning
Ziqi Zhao, Zhaochun Ren, Liu Yang +6
Offline reinforcement learning (RL) aims to learn policies without online explorations. To enlarge the training data, model-based offline RL learns a dynamics model which is utiliz…
Improving Sequential Recommenders through Counterfactual Augmentation of System Exposure
Ziqi Zhao, Zhaochun Ren, Jiyuan Yang +7
In sequential recommendation (SR), system exposure refers to items that are exposed to the user. Typically, only a few of the exposed items would be interacted with by the user. Al…