2 papers
cs.DC2025
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
Bowen Pang, Kai Li, Feifan Wang
The increasing adoption of large language models (LLMs) necessitates inference serving systems that can deliver both high throughput and low latency. Deploying LLMs with hundreds o…
cs.DC2025
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization
Bowen Pang, Kai Li, Ruifeng She +1
With the development of large language models (LLMs), it has become increasingly important to optimize hardware usage and improve throughput. In this paper, we study the inference…