5 papers
FMplex: Model Virtualization for Serving Extensible Foundation Models
Hetvi Shastri, Pragya Sharma, Walid A. Hanafy +3
Foundation models (FMs) are increasingly used as backbones for downstream tasks across language, vision, time-series, and multimodal applications. Yet existing model-serving system…
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
Nathan Ng, Walid A. Hanafy, Prashanthi Kadambi +5
IoT applications increasingly rely on on-device AI accelerators to ensure high performance, especially in low-connectivity and safety-critical scenarios. However, the limited on-ch…
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
Nathan Ng, David Irwin, Ananthram Swami +2
Computational offloading is a promising approach for overcoming resource constraints on client devices by moving some or all of an application's computations to remote servers. Wit…
FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
Li Wu, Walid A. Hanafy, Tarek Abdelzaher +3
Model serving systems have become popular for deploying deep learning models for various latency-sensitive inference tasks. While traditional replication-based methods have been us…
LLM-Driven Auto Configuration for Transient IoT Device Collaboration
Hetvi Shastri, Walid A. Hanafy, Li Wu +3
Today's Internet of Things (IoT) has evolved from simple sensing and actuation devices to those with embedded processing and intelligent services, enabling rich collaborations betw…