4 citations · 10 across the 5 of their papers we have counts for
5 papers
Automatic Attention Pruning: Improving and Automating Model Pruning using Attentions
Kaiqi Zhao, Animesh Jain, Ming Zhao
Pruning is a promising approach to compress deep learning models in order to deploy them on resource-constrained edge devices. However, many existing pruning solutions are based on…
A Contrastive Knowledge Transfer Framework for Model Compression and Transfer Learning
Kaiqi Zhao, Yitao Chen, Ming Zhao
Knowledge Transfer (KT) achieves competitive performance and is widely used for image classification tasks in model compression and transfer learning. Existing KT works transfer th…
GPU-enabled Function-as-a-Service for Machine Learning Inference
Ming Zhao, Kritshekhar Jha, Sungho Hong
Function-as-a-Service (FaaS) is emerging as an important cloud computing service model as it can improve the scalability and usability of a wide range of applications, especially M…
Enabling Deep Learning on Edge Devices through Filter Pruning and Knowledge Transfer
Kaiqi Zhao, Yitao Chen, Ming Zhao
Deep learning models have introduced various intelligent applications to edge devices, such as image classification, speech recognition, and augmented reality. There is an increasi…
Iterative Activation-based Structured Pruning
Kaiqi Zhao, Animesh Jain, Ming Zhao
Deploying complex deep learning models on edge devices is challenging because they have substantial compute and memory resource requirements, whereas edge devices' resource budget…