papers
Publications (3)
cs.DC2025
WindVE: Collaborative CPU-NPU Vector Embedding
Jinqi Huang, Xuebing Yu, Yi Xiong +4
Retrieval-Augmented Generation is a technology that enhances large language models by integrating information retrieval. In the industry, inference services based on LLMs are highl…
cs.DC2025
High-Throughput LLM inference on Heterogeneous Clusters
Yi Xiong, Jinqi Huang, Wenjie Huang +6
Nowadays, many companies possess various types of AI accelerators, forming heterogeneous clusters. Efficiently leveraging these clusters for high-throughput large language model (L…
cs.DC2025
SLO-Aware Scheduling for Large Language Model Inferences
Jinqi Huang, Yi Xiong, Xuebing Yu +4
Large language models (LLMs) have revolutionized applications such as code completion, chatbots, and online classification. To elevate user experiences, service level objectives (S…