3 citations · 3 across the 2 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2025
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
Siyuan Chen, Zhipeng Jia, Samira Khan +2
This paper introduces SLOs-Serve, a system designed for serving multi-stage large language model (LLM) requests with application- and stage-specific service level objectives (SLOs)…
cs.DC2024
Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors
Siyuan Chen, Zhuofeng Wang, Zelong Guan +2
Fine-tuning large language models (LLMs) requires significant memory, often exceeding the capacity of a single GPU. A common solution to this memory challenge is offloading compute…