4 papers
TURBO: Utility-Aware Bandwidth Allocation for Cloud-Augmented Autonomous Control
Peter Schafhalter, Alexander Krentsel, Hongbo Wei +4
Autonomous driving system progress has been driven by improvements in machine learning models, whose computational demands now exceed what edge devices alone can provide. The cloud…
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
Shiyi Cao, Shu Liu, Tyler Griggs +6
Efficient deployment of large language models, particularly Mixture of Experts (MoE), on resource-constrained platforms presents significant challenges, especially in terms of comp…
Scalable Multi-Domain Adaptation of Language Models using Modular Experts
Peter Schafhalter, Shun Liao, Yanqi Zhou +3
Domain-specific adaptation is critical to maximizing the performance of pre-trained language models (PLMs) on one or multiple targeted tasks, especially under resource-constrained…
Managing Bandwidth: The Key to Cloud-Assisted Autonomous Driving
Alexander Krentsel, Peter Schafhalter, Joseph E. Gonzalez +3
Prevailing wisdom asserts that one cannot rely on the cloud for critical real-time control systems like self-driving cars. We argue that we can, and must. Following the trends of i…