3 papers
cs.LG2022
Demand Layering for Real-Time DNN Inference with Minimized Memory Usage
Mingoo Ji, Saehanseul Yi, Changjin Koo +4
When executing a deep neural network (DNN), its model parameters are loaded into GPU memory before execution, incurring a significant GPU memory burden. There are studies that redu…
cs.LG2022
Hybrid Learning for Orchestrating Deep Learning Inference in Multi-user Edge-cloud Networks
Sina Shahhosseini, Tianyi Hu, Dongjoo Seo +4
Deep-learning-based intelligent services have become prevalent in cyber-physical applications including smart cities and health-care. Collaborative end-edge-cloud computing for dee…
cs.LG2022
Online Learning for Orchestration of Inference in Multi-User End-Edge-Cloud Networks
Sina Shahhosseini, Dongjoo Seo, Anil Kanduri +5
Deep-learning-based intelligent services have become prevalent in cyber-physical applications including smart cities and health-care. Deploying deep-learning-based intelligence nea…