12 papers
Kwai Summary Attention Technical Report
Chenglong Chu, Guorui Zhou, Guowang Zhang +35
Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agen…
Quantized Inference for OneRec-V2
Yi Su, Xinchen Luo, Hongtao Cheng +7
Quantized inference has demonstrated substantial system-level benefits in large language models while preserving model quality. In contrast, reliably applying low-precision quantiz…
Kelix Technical Report
Boyang Ding, Chenglong Chu, Dunju Zang +28
Autoregressive large language models (LLMs) scale well by expressing diverse tasks as sequences of discrete natural-language tokens and training with next-token prediction, which u…
OneLive: Dynamically Unified Generative Framework for Live-Streaming Recommendation
Shen Wang, Yusheng Huang, Ruochen Yang +15
Live-streaming recommender system serves as critical infrastructure that bridges the patterns of real-time interactions between users and authors. Similar to traditional industrial…
PIT: A Dynamic Personalized Item Tokenizer for End-to-End Generative Recommendation
Huanjie Wang, Xinchen Luo, Honghui Bao +6
Generative Recommendation has revolutionized recommender systems by reformulating retrieval as a sequence generation task over discrete item identifiers. Despite the progress, exis…
OpenOneRec Technical Report
Guorui Zhou, Honghui Bao, Jiaming Huang +44
While the OneRec series has successfully unified the fragmented recommendation pipeline into an end-to-end generative framework, a significant gap remains between recommendation sy…