2 citations · 3 across the 7 of their papers we have counts for
15 papers
Teacher Retains Full Tokens, Student Merges Efficiently: TM20K for E-Commerce Sequence Modeling in Ad Recommendation
Xinchun Li, Duoru Zheng, Wenlin Zhao +13
Benefiting from ultra-long behavior sequence modeling, existing recommender systems bring users a better experience via simultaneously considering their long-term and short-term in…
Rec-Distill: An Industrial Distillation Pipeline for Large-Scale Recommendation Models
Haoran Ding, Wenlin Zhao, Yuchen Jiang +16
Large recommendation models have demonstrated substantial potential gains under scaling laws, yet these gains are difficult to realize in industrial recommendation systems because…
Bottleneck Tokens for Unified Multimodal Retrieval
Siyu Sun, Jing Ren, Zhaohe Liao +8
Adapting decoder-only multimodal large language models (MLLMs) for unified multimodal retrieval faces two structural gaps. First, existing methods rely on implicit pooling, which o…
TokenMixer-Large: Scaling Up Large Ranking Models in Industrial Recommenders
Yuchen Jiang, Jie Zhu, Xintian Han +18
While scaling laws for recommendation models have gained significant traction, existing architectures such as Wukong, HiFormer and DHEN, often struggle with sub-optimal designs and…
MSN: A Memory-based Sparse Activation Scaling Framework for Large-scale Industrial Recommendation
Shikang Wu, Hui Lu, Jinqiu Jin +9
Scaling deep learning recommendation models is an effective way to improve model expressiveness. Existing approaches often incur substantial computational overhead, making them dif…
MixFormer: Co-Scaling Up Dense and Sequence in Industrial Recommenders
Xu Huang, Hao Zhang, Zhifang Fan +6
As industrial recommender systems enter a scaling-driven regime, Transformer architectures have become increasingly attractive for scaling models towards larger capacity and longer…