2 papers
cs.DC2025
CPU-Limits kill Performance: Time to rethink Resource Control
Chirag Shetty, Sarthak Chakraborty, Hubertus Franke +4
Research in compute resource management for cloud-native applications is dominated by the problem of setting optimal CPU limits -- a fundamental OS mechanism that strictly restrict…
cs.DC2024
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
Haoran Qiu, Weichao Mao, Archit Patke +7
Large language models (LLMs) have been driving a new wave of interactive AI applications across numerous domains. However, efficiently serving LLM inference requests is challenging…