3 citations · 3 across the 2 of their papers we have counts for
3 papers
cs.DC2026
SMART-MIG: A Learning Framework for Scalable and Energy-Efficient GPU Scheduling
Wenqing Yu, Neel Karia, Tanvi Hisaria +3
The emergence of Multi-Instance GPU (MIG) technology enables us to run smaller machine learning models on partitions of a GPU rather than the entire device, thus improving utilizat…
cs.DC2026★ 3 cited
Energy Efficient Scheduling of AI/ML Workloads on Multi Instance GPUs with Dynamic Repartitioning
Ellie Lipe, Neel Karia, Connor Espenshade +3
Increasing demand from AI/ML workloads is exacerbating the rising energy consumption of data centers. Recent advances in hardware such as NVIDIA's Multi Instance GPUs (MIGs) offer…
cs.ET2026
WVA: A Global Optimization Control Plane for llmd
Abhishek Malvankar, Lionel Villard, Mohammed Abdi +6
As Large Language Models (LLMs) scale to handle massive concurrent traffic, optimizing the infrastructure required for inference has become a primary challenge. To manage the high…