Serve Programs, Not Prompts
arXiv:2510.25412 · doi:10.1145/3713082.3730398
Abstract
Current large language model (LLM) serving systems, primarily designed for text completion, are neither efficient nor adaptable for increasingly complex LLM applications due to their inflexible design. We propose a new LLM serving system architecture that serves programs instead of prompts to address this problem. These programs, called LLM Inference Programs (LIPs), allow users to customize token prediction and KV cache management at runtime and to offload parts of their application logic, such as tool execution, to the server. We describe an example of this architecture through a system named Symphony, which functions as an operating system for LIPs. Symphony exposes LLM model computations via system calls and virtualizes KV cache with a dedicated file system, while ensuring GPU efficiency with a two-level process scheduling scheme. Symphony has the potential to open the door to a more efficient and extensible ecosystem for LLM applications.
HotOS 2025. Follow-up implementation work (SOSP 2025) is available at arXiv:2510.24051
References in corpus (16)
- ReAct: Synergizing Reasoning and Acting in Language Models
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Prompting Is Programming: A Query Language for Large Language Models
- Gorilla: Large Language Model Connected with Massive APIs
- Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
- Efficient Streaming Language Models with Attention Sinks
- Fast Inference from Transformers via Speculative Decoding
- HO: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
- Prompt Cache: Modular Attention Reuse for Low-Latency Inference
- XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
- Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
- EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
- Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation
- Language Model Cascades: Token-level uncertainty and beyond