4 papers
Generation Quality-Latency Tradeoff-Aware Inference Offloading for Multimodal LLMs in Cloud-Edge Continuum
Zhongxiao Wang, Yueshen Xu, Xinkui Zhao +2
Beyond pure cloud, some efforts are being made to deploy Large Language Models (LLMs) in edge to accelerate inference response. So the deployment of LLMs in cloud-edge continuum be…
How Well Does Generative Recommendation Generalize?
Yijie Ding, Zitian Guo, Jiacheng Li +8
A widely held hypothesis for why generative recommendation (GR) models outperform conventional item ID-based models is that they generalize better. However, there is few systematic…
Preference Discerning with LLM-Enhanced Generative Retrieval
Fabian Paischer, Liu Yang, Linfeng Liu +12
In sequential recommendation, models recommend items based on user's interaction history. To this end, current models usually incorporate information such as item descriptions and…
Generating Long Semantic IDs in Parallel for Recommendation
Yupeng Hou, Jiacheng Li, Ashley Shin +6
Semantic ID-based recommendation models tokenize each item into a small number of discrete tokens that preserve specific semantics, leading to better performance, scalability, and…