2 citations · 3 across the 4 of their papers we have counts for
4 papers
CES-KD: Curriculum-based Expert Selection for Guided Knowledge Distillation
Ibtihel Amara, Maryam Ziaeefard, Brett H. Meyer +2
Knowledge distillation (KD) is an effective tool for compressing deep classification models for edge devices. However, the performance of KD is affected by the large capacity gap b…
Consistency driven Sequential Transformers Attention Model for Partially Observable Scenes
Samrudhdhi B. Rangrej, Chetan L. Srinidhi, James J. Clark
Most hard attention models initially observe a complete scene to locate and sense informative glimpses, and predict class-label of a scene based on glimpses. However, in many appli…
Standard Deviation-Based Quantization for Deep Neural Networks
Amir Ardakani, Arash Ardakani, Brett Meyer +2
Quantization of deep neural networks is a promising approach that reduces the inference cost, making it feasible to run deep networks on resource-restricted devices. Inspired by ex…
Kronecker Decomposition for GPT Compression
Ali Edalati, Marzieh Tahaei, Ahmad Rashid +3
GPT is an auto-regressive Transformer-based pre-trained language model which has attracted a lot of attention in the natural language processing (NLP) domain due to its state-of-th…