2 papers
cs.LG2026
Make It Long, Keep It Fast: End-to-End 10K Long User Behavior Sequence Modeling for Billion-Scale Douyin Recommendation
Lin Guan, Jia-Qi Yang, Zhishan Zhao +12
Short-video recommenders such as Douyin must exploit extremely long user behavior histories without breaking latency or cost budgets. We present an end-to-end industrial recommende…
cs.LG2025
AIF: Asynchronous Inference Framework for Cost-Effective Pre-Ranking
Zhi Kou, Xiang-Rong Sheng, Shuguang Han +5
In industrial recommendation systems, pre-ranking models based on deep neural networks (DNNs) commonly adopt a sequential execution framework: feature fetching and model forward co…