5 citations · 6 across the 5 of their papers we have counts for
5 papers
Energy-Aware Scheduling for Serverless LLM Serving on Shared GPUs
Tianyu Wang, Gourav Rattihalli, Aditya Dhakal +2
As LLM inference becomes a major cloud workload, its growing energy footprint makes cluster-wide energy optimization increasingly important. Serverless LLM serving helps platforms…
Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding
Tianyu Wang, Gourav Rattihalli, Aditya Dhakal +4
Dynamic sparse attention (DSA) accelerates long-context LLM decoding by attending to only the top-K KV blocks relevant to each query, but it introduces a serialized selection-to-at…
An RDMA-First Object Storage System with SmartNIC Offload
Yu Zhu, Aditya Dhakal, Pedro Bruel +4
AI training and inference impose sustained, fine-grain I/O that stresses host-mediated, TCP-based storage paths. Motivated by kernel-bypass networking and user-space storage stacks…
GreenFaaS: Maximizing Energy Efficiency of HPC Workloads with FaaS
Alok Kamatar, Valerie Hayot-Sasson, Yadu Babuji +6
Application energy efficiency can be improved by executing each application component on the compute element that consumes the least energy while also satisfying time constraints.…
Two stage cluster for resource optimization with Apache Mesos
Gourav Rattihalli, Pankaj Saha, Madhusudhan Govindaraju +1
As resource estimation for jobs is difficult, users often overestimate their requirements. Both commercial clouds and academic campus clusters suffer from low resource utilization…