activity
20242026
collaborators

11 papers

cs.LG2026

TWICE: Two-Clock, Two-Window Learning for Long-Horizon Conversion Prediction in Online Advertising

Kaiyuan Li, Kun Wang, Zhongbo Wang +4

Long-horizon conversion prediction under delayed feedback creates a two-clock, two-window learning problem in online advertising. A short base observation window releases recent cl…

cs.LG2026

ACPO: Asymmetric Credit Policy Optimization via Mode-Local Entropy Surrogate

Zijun Xie, Yuyang You, Yongzhi Li +8

Outcome-supervised reinforcement learning scales to verifiable reasoning tasks, but trajectory-level rewards assign the same outcome signal to all sampled tokens, overlooking their…

cs.LG2026

FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning

Xikai Zhang, Yongzhi Li, Likang Xiao +6

Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, the training loop of GRPO and its variants…

cs.IR2026

Aggregate and Broadcast: Scalable and Efficient Feature Interaction for Recommender Systems

Kaiyuan Li, Yongxiang Tang, Wenzheng Shu +5

Feature interaction is a core ingredient in ranking models for large-scale recommender systems, yet making it both expressive and efficiently scalable remains challenging. Exhausti…

cs.IR2025

VQL: An End-to-End Context-Aware Vector Quantization Attention for Ultra-Long User Behavior Modeling

Kaiyuan Li, Yongxiang Tang, Yanhua Cheng +5

In large-scale recommender systems, ultra-long user behavior sequences encode rich signals of evolving interests. Extending sequence length generally improves accuracy, but directl…

cs.IR2025

Reward Balancing Revisited: Enhancing Offline Reinforcement Learning for Recommender Systems

Wenzheng Shu, Yanxiang Zeng, Yongxiang Tang +6

Offline reinforcement learning (RL) has emerged as a prevalent and effective methodology for real-world recommender systems, enabling learning policies from historical data and cap…