Showing cs.OSShow all
2 papers · 1 filter
cs.OS2026
Adaptive Context Parallelism for Production LLM Serving
Jiarui Guo, Rongle Wang, Peijun Huang +7
As LLM context windows expand and input sequences grow longer, serving systems face increasing computational and memory demands. Context parallelism (CP), which partitions the inpu…
cs.OS2026
RTP-LLM: High-Performance Alibaba LLM Inference Engine
Boyu Tan, Jiarui Guo, Zongwei Lv +26
Large Language Models (LLMs) have revolutionized AI applications, but deploying them at scale presents significant challenges. We present RTP-LLM, a high-performance inference engi…