76 citations · 218 across the 8 of their papers we have counts for
8 papers
FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic Rounding
Sai Qian Zhang, Bradley McDanel, H. T. Kung
Block Floating Point (BFP) can efficiently support quantization for Deep Neural Network (DNN) training by providing a wide dynamic range via a shared exponent across a group of val…
Term Revealing: Furthering Quantization at Run Time on Quantized DNNs
H. T. Kung, Bradley McDanel, Sai Qian Zhang
We present a novel technique, called Term Revealing (TR), for furthering quantization at run time for improved performance of Deep Neural Networks (DNNs) already quantized with con…
Full-stack Optimization for Accelerating CNNs with FPGA Validation
Bradley McDanel, Sai Qian Zhang, H. T. Kung +1
We present a full-stack optimization framework for accelerating inference of CNNs (Convolutional Neural Networks) and validate the approach with field-programmable gate arrays (FPG…
Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint Optimization
H. T. Kung, Bradley McDanel, Sai Qian Zhang
This paper describes a novel approach of packing sparse convolutional neural networks for their efficient systolic array implementations. By combining subsets of columns in the ori…
Incomplete Dot Products for Dynamic Computation Scaling in Neural Network Inference
Bradley McDanel, Surat Teerapittayanon, H. T. Kung
We propose the use of incomplete dot products (IDP) to dynamically adjust the number of input channels used in each layer of a convolutional neural network during feedforward infer…
Embedded Binarized Neural Networks
Bradley McDanel, Surat Teerapittayanon, H. T. Kung
We study embedded Binarized Neural Networks (eBNNs) with the aim of allowing current binarized neural networks (BNNs) in the literature to perform feedforward inference efficiently…