4 papers
Cross-Scenario Unified Modeling of User Interests at Billion Scale
Manjie Xu, Cheng Chen, Xin Jia +9
User interests on content platforms are inherently diverse, manifesting through complex behavioral patterns across heterogeneous scenarios such as search, feed browsing, and conten…
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
Yifan Qiao, Shu Anzai, Shan Yu +10
Large language model (LLM) serving demands low latency and high throughput, but high load variability makes it challenging to achieve high GPU utilization. In this paper, we identi…
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
Chaoyi Ruan, Yinhe Chen, Dongqi Tian +4
LLM inference must meet strict latency SLOs (e.g., 100 ms P99 time-between-tokens) while maximizing goodput. Yet, real-world variability in prompt and response lengths skews comput…
Prompt Inversion Attack against Collaborative Inference of Large Language Models
Wenjie Qu, Yuguang Zhou, Yongji Wu +4
Large language models (LLMs) have been widely applied for their remarkable capability of content generation. However, the practical use of open-source LLMs is hindered by high reso…