2 papers
cs.DC2024
Optimal Workload Placement on Multi-Instance GPUs
Bekir Turkkan, Pavankumar Murali, Pavithra Harsha +3
There is an urgent and pressing need to optimize usage of Graphical Processing Units (GPUs), which have arguably become one of the most expensive and sought after IT resources. To…
cs.LG2024
Leveraging Interpretability in the Transformer to Automate the Proactive Scaling of Cloud Resources
Amadou Ba, Pavithra Harsha, Chitra Subramanian
Modern web services adopt cloud-native principles to leverage the advantages of microservices. To consistently guarantee high Quality of Service (QoS) according to Service Level Ag…