5 papers
Mitigating context switching in densely packed Linux clusters with Latency-Aware Group Scheduling
Al Amjad Tawfiq Isstaif, Evangelia Kalyvianaki, Richard Mortier
Cluster orchestrators such as Kubernetes depend on accurate estimates of node capacity and job requirements. Inaccuracies in either lead to poor placement decisions and degraded cl…
LSKV: A Confidential Distributed Datastore to Protect Critical Data in the Cloud
Andrew Jeffery, Julien Maffre, Heidi Howard +1
Software services are increasingly migrating to the cloud, requiring trust in actors with direct access to the hardware, software and data comprising the service. A distributed dat…
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
Andrew Jeffery, Chris Jensen, Richard Mortier
Application tail latency is a key metric for many services, with high latencies being linked directly to loss of revenue. Modern deeply-nested micro-service architectures exacerbat…
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
Grant Wilkins, Srinivasan Keshav, Richard Mortier
The rapid adoption of large language models (LLMs) has led to significant advances in natural language processing and text generation. However, the energy consumed through LLM mode…
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
Grant Wilkins, Srinivasan Keshav, Richard Mortier
Both the training and use of Large Language Models (LLMs) require large amounts of energy. Their increasing popularity, therefore, raises critical concerns regarding the energy eff…