3 papers
cs.IR2026
Towards Faithful Simulation of Human Shopping Behavior
Jiakai Tang, Yan Mi, Jing Yu +9
Simulating realistic user shopping behavior underpins offline evaluation and reinforcement learning in e-commerce scenarios. While recent LLM- and VLM-based simulators have made en…
cs.LG2026
Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics
Longtian Bao, Jianyou Wang, Yang Zhang +2
Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning traces are usually unavailable,…
cs.IR2026
Learning from the Future: Privileged Self-Distillation for Sequential Recommendation
Jiakai Tang, Yang Zhang, See-Kiong Ng +4
Sequential recommenders are commonly trained with one-hot next-item labels under a causal (prefix-only) objective aligned with inference. While deployment-compatible, this supervis…