128 citations · 128 across the 8 of their papers we have counts for
8 papers
Off-Policy Primal-Dual Safe Reinforcement Learning
Zifan Wu, Bo Tang, Qian Lin +5
Primal-dual safe RL methods commonly perform iterations between the primal update of the policy and the dual update of the Lagrange Multiplier. Such a training paradigm is highly s…
Deep Automated Mechanism Design for Integrating Ad Auction and Allocation in Feed
Xuejian Li, Ze Wang, Bingqi Zhu +3
E-commerce platforms usually present an ordered list, mixed with several organic items and an advertisement, in response to each user's page view request. This list, the outcome of…
TBIN: Modeling Long Textual Behavior Data for CTR Prediction
Shuwei Chen, Xiang Li, Jian Dong +3
Click-through rate (CTR) prediction plays a pivotal role in the success of recommendations. Inspired by the recent thriving of language models (LMs), a surge of works improve predi…
MDDL: A Framework for Reinforcement Learning-based Position Allocation in Multi-Channel Feed
Xiaowen Shi, Ze Wang, Yuanying Cai +6
Nowadays, the mainstream approach in position allocation system is to utilize a reinforcement learning model to allocate appropriate locations for items in various channels and the…
PIER: Permutation-Level Interest-Based End-to-End Re-ranking Framework in E-commerce
Xiaowen Shi, Fan Yang, Ze Wang +6
Re-ranking draws increased attention on both academics and industries, which rearranges the ranking list by modeling the mutual influence among items to better meet users' demands.…
A Deep Behavior Path Matching Network for Click-Through Rate Prediction
Jian Dong, Yisong Yu, Yapeng Zhang +6
User behaviors on an e-commerce app not only contain different kinds of feedback on items but also sometimes imply the cognitive clue of the user's decision-making. For understandi…