activity
20172022
most citedPatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight Pruning

211 citations · 871 across the 52 of their papers we have counts for

collaborators

97 papers

cs.LG202228 cited

VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformer

Mengshu Sun, Haoyu Ma, Guoliang Kang +5

The transformer architectures with attention mechanisms have obtained success in Nature Language Processing (NLP), and Vision Transformers (ViTs) have recently extended the applica…

cs.CV202221 cited

F8Net: Fixed-Point 8-bit Only Multiplication for Network Quantization

Qing Jin, Jian Ren, Richard Zhuang +6

Neural network quantization is a promising compression technique to reduce memory footprint and save energy consumption, potentially leading to real-time inference. However, there…

cs.CV202113 cited

ScaleCert: Scalable Certified Defense against Adversarial Patches with Sparse Superficial Layers

Husheng Han, Kaidi Xu, Xing Hu +6

Adversarial patch attacks that craft the pixels in a confined region of the input images show their powerful attack effectiveness in physical environments even with noises or defor…

cs.LG20211 cited

ILMPQ : An Intra-Layer Multi-Precision Deep Neural Network Quantization framework for FPGA

Sung-En Chang, Yanyu Li, Mengshu Sun +2

This work targets the commonly used FPGA (field-programmable gate array) devices as the hardware platform for DNN edge computing. We focus on DNN quantization as the main model com…

cs.LG2021

RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple Precisions

Sung-En Chang, Yanyu Li, Mengshu Sun +4

This work proposes a novel Deep Neural Network (DNN) quantization framework, namely RMSMP, with a Row-wise Mixed-Scheme and Multi-Precision approach. Specifically, this is the firs…

cs.LG202141 cited

MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the Edge

Geng Yuan, Xiaolong Ma, Wei Niu +13

Recently, a new trend of exploring sparsity for accelerating neural network training has emerged, embracing the paradigm of training on the edge. This paper proposes a novel Memory…