2 papers
cs.OS2026
Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate
Fangyue Liu, Hua Liu, Xinyuan Lyu +8
LLM inference powers latency-critical production services nowadays. The bursty nature of inference traffic results in over-provisioning, which in turn leads to resource underutiliz…
cs.AR2026
ROMA: a Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
Wenqiang Wang, Yijia Zhang, Zikai Zhang +4
As large language models (LLMs) demonstrate powerful capabilities, deploying them on edge devices has become increasingly crucial, offering advantages in privacy and real-time inte…