4 papers · 1 filter
WindVE: Collaborative CPU-NPU Vector Embedding
Jinqi Huang, Xuebing Yu, Yi Xiong +4
Retrieval-Augmented Generation is a technology that enhances large language models by integrating information retrieval. In the industry, inference services based on LLMs are highl…
SLO-Aware Scheduling for Large Language Model Inferences
Jinqi Huang, Yi Xiong, Xuebing Yu +4
Large language models (LLMs) have revolutionized applications such as code completion, chatbots, and online classification. To elevate user experiences, service level objectives (S…
High-Throughput LLM inference on Heterogeneous Clusters
Yi Xiong, Jinqi Huang, Wenjie Huang +6
Nowadays, many companies possess various types of AI accelerators, forming heterogeneous clusters. Efficiently leveraging these clusters for high-throughput large language model (L…
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
Ping Zhang, Lei Su, Jinjie Yang +1
Hosting diverse large language model workloads in a unified resource pool through co-location is cost-effective. For example, long-running chat services generally follow diurnal tr…