5 papers
A Performance Analyzer for a Public Cloud's ML-Augmented VM Allocator
Roozbeh Bostandoost, Pooria Namyar, Siva Kesava Reddy Kakarla +8
Cloud operators increasingly deploy multiple ML models in their VM allocation pipelines. In such settings, individually benign predictions can shift and compound, severely degradin…
Harvest: Adaptive Photonic Switching Schedules for Collective Communication in Scale-up Domains
Mahir Rahman, Samuel Joseph, Nihar Kodkani +2
As chip-to-chip silicon photonics gain traction for their bandwidth and energy efficiency, their circuit-switched nature raises a fundamental question for collective communication:…
Dynamic Rebatching for Efficient Early-Exit Inference with DREX
Xuting Liu, Daniel Alexander, Siva Kesava Reddy Kakarla +2
Early-Exit (EE) is a Large Language Model (LLM) architecture that accelerates inference by allowing easier tokens to be generated using only a subset of the model's layers. However…
Robust Heuristic Algorithm Design with LLMs
Pantea Karimi, Dany Rouhana, Pooria Namyar +3
We posit that we can generate more robust and performant heuristics if we augment approaches using LLMs for heuristic design with tools that explain why heuristics underperform and…
Enhancing Network Failure Mitigation with Performance-Aware Ranking
Pooria Namyar, Arvin Ghavidel, Daniel Crankshaw +5
Cloud providers install mitigations to reduce the impact of network failures within their datacenters. Existing network mitigation systems rely on simple local criteria or global p…