7 papers
RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
Fanqing Meng, Lingxiao Du, Qiguang Chen +4
Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capability gaps, designing and val…
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
Ziqi Zhao, Xinyu Ma, Liu Yang +6
On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-policy rollouts. However, ex…
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Fanqing Meng, Lingxiao Du, Zijian Wu +46
Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surrounding environment may change in…
Curriculum Approximate Unlearning for Session-based Recommendation
Liu Yang, Zhaochun Ren, Ziqi Zhao +7
Approximate unlearning for session-based recommendation refers to eliminating the influence of specific training samples from the recommender without retraining of (sub-)models. Gr…
Offline Trajectory Optimization for Offline Reinforcement Learning
Ziqi Zhao, Zhaochun Ren, Liu Yang +6
Offline reinforcement learning (RL) aims to learn policies without online explorations. To enlarge the training data, model-based offline RL learns a dynamics model which is utiliz…
Improving Sequential Recommenders through Counterfactual Augmentation of System Exposure
Ziqi Zhao, Zhaochun Ren, Jiyuan Yang +7
In sequential recommendation (SR), system exposure refers to items that are exposed to the user. Typically, only a few of the exposed items would be interacted with by the user. Al…