Showing cs.DCShow all
2 papers · 1 filter
cs.DC2024
Teola: Towards End-to-End Optimization of LLM-based Applications
Xin Tan, Yimin Jiang, Yitao Yang +1
Large language model (LLM)-based applications consist of both LLM and non-LLM components, each contributing to the end-to-end latency. Despite great efforts to optimize LLM inferen…
cs.DC2024
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
Bodun Hu, Jiamin Li, Le Xu +5
The increasing demand for Large Language Models (LLMs) across various applications has led to a significant shift in the design of deep learning serving systems. Deploying LLMs, pa…