Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
arXiv:1510.00149
Abstract
Neural networks are both computationally intensive and memory intensive, making them difficult to deploy on embedded systems with limited hardware resources. To address this limitation, we introduce "deep compression", a three stage pipeline: pruning, trained quantization and Huffman coding, that work together to reduce the storage requirement of neural networks by 35x to 49x without affecting their accuracy. Our method first prunes the network by learning only the important connections. Next, we quantize the weights to enforce weight sharing, finally, we apply Huffman coding. After the first two steps we retrain the network to fine tune the remaining connections and the quantized centroids. Pruning, reduces the number of connections by 9x to 13x; Quantization then reduces the number of bits that represent each connection from 32 to 5. On the ImageNet dataset, our method reduced the storage required by AlexNet by 35x, from 240MB to 6.9MB, without loss of accuracy. Our method reduced the size of VGG-16 by 49x from 552MB to 11.3MB, again with no loss of accuracy. This allows fitting the model into on-chip SRAM cache rather than off-chip DRAM memory. Our compression method also facilitates the use of complex neural networks in mobile applications where application size and download bandwidth are constrained. Benchmarked on CPU, GPU and mobile GPU, compressed network has 3x to 4x layerwise speedup and 3x to 7x better energy efficiency.
Published as a conference paper at ICLR 2016 (oral)
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- Compressing Deep Convolutional Networks using Vector Quantization
- Compressing Neural Networks with the Hashing Trick
- Memory Bounded Deep Convolutional Networks
Cited by in corpus (314)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
- Deep Multi-modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges
- Deep Face Recognition: A Survey
- Benchmark Analysis of Representative Deep Neural Network Architectures
- Trained Ternary Quantization
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- FPGA-based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- A Systematic DNN Weight Pruning Framework using Alternating Direction Method of Multipliers
- Q8BERT: Quantized 8Bit BERT
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
- Fast-SCNN: Fast Semantic Segmentation Network
- Channel Pruning for Accelerating Very Deep Neural Networks
- SpArch: Efficient Architecture for Sparse Matrix Multiplication
- Model compression via distillation and quantization
- Machine Learning for Microcontroller-Class Hardware: A Review
- Radar-Camera Fusion for Object Detection and Semantic Segmentation in Autonomous Driving: A Comprehensive Review
- A Microprocessor implemented in 65nm CMOS with Configurable and Bit-scalable Accelerator for Programmable In-memory Computing
- Learning New Physics from a Machine
- Driver Drowsiness Detection Model Using Convolutional Neural Networks Techniques for Android Application
- GhostNets on Heterogeneous Devices via Cheap Operations
- Empowering Edge Intelligence: A Comprehensive Survey on On-Device AI Models
- QuantumNAS: Noise-Adaptive Search for Robust Quantum Circuits
- CirCNN: Accelerating and Compressing Deep Neural Networks Using Block-CirculantWeight Matrices
- Tiny Machine Learning: Progress and Futures
- The NLP Cookbook: Modern Recipes for Transformer based Deep Learning Architectures
- JALAD: Joint Accuracy- and Latency-Aware Deep Structure Decoupling for Edge-Cloud Execution
- Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy
- Distilled Siamese Networks for Visual Tracking
- Automatic Sleep Staging of EEG Signals: Recent Development, Challenges, and Future Directions
- FETCH: A deep-learning based classifier for fast transient classification
- Exploring Sparsity in Recurrent Neural Networks
- Faster CNNs with Direct Sparse Convolutions and Guided Pruning
- SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks
- A Theoretical Analysis of Deep Q-Learning
- Learning Intrinsic Sparse Structures within Long Short-Term Memory
- Reinventing 2D Convolutions for 3D Images
- Compression of Deep Convolutional Neural Networks for Fast and Low Power Mobile Applications
- Light-Weight RefineNet for Real-Time Semantic Segmentation
- Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights
- A Survey on Approximate Edge AI for Energy Efficient Autonomous Driving Services
- Deep Learning Methods for Fingerprint-Based Indoor Positioning: A Review
- Layer-specific Optimization for Mixed Data Flow with Mixed Precision in FPGA Design for CNN-based Object Detectors
- Accelerator-Aware Pruning for Convolutional Neural Networks
- Dual Dynamic Inference: Enabling More Efficient, Adaptive and Controllable Deep Inference
- A TinyML Platform for On-Device Continual Learning with Quantized Latent Replays
- Deep Learning Methods for Solving Linear Inverse Problems: Research Directions and Paradigms
- Robust Ultra-wideband Range Error Mitigation with Deep Learning at the Edge
- A Low Effort Approach to Structured CNN Design Using PCA
- The Global Landscape of Neural Networks: An Overview
- FSpiNN: An Optimization Framework for Memory- and Energy-Efficient Spiking Neural Networks
- The unreasonable effectiveness of the forget gate
- Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks
- AddNet: Deep Neural Networks Using FPGA-Optimized Multipliers
- Federated Learning for Healthcare Informatics
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- Shape-based Magnetic Domain Wall Drift for an Artificial Spintronic Leaky Integrate-and-Fire Neuron
- Filter Pruning by Switching to Neighboring CNNs with Good Attributes
- Integrated Photonic Tensor Processing Unit for a Matrix Multiply: a Review
- EdgeDRNN: Recurrent Neural Network Accelerator for Edge Inference
- Feature Pyramid and Hierarchical Boosting Network for Pavement Crack Detection
- Sobolev Training for Neural Networks
- Real-Time Decoding for Fault-Tolerant Quantum Computing: Progress, Challenges and Outlook
- EDropout: Energy-Based Dropout and Pruning of Deep Neural Networks
- MaskConnect: Connectivity Learning by Gradient Descent
- Sparsity through evolutionary pruning prevents neuronal networks from overfitting
- Cluster Pruning: An Efficient Filter Pruning Method for Edge AI Vision Applications
- Reducing Computational Complexity of Neural Networks in Optical Channel Equalization: From Concepts to Implementation
- Insights on representational similarity in neural networks with canonical correlation
- Edge Deep Learning for Neural Implants
- FermiNets: Learning generative machines to generate efficient neural networks via generative synthesis
- Lets keep it simple, Using simple architectures to outperform deeper and more complex architectures
- Lightweight Modules for Efficient Deep Learning based Image Restoration
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Automated Pruning for Deep Neural Network Compression
- PruneTrain: Fast Neural Network Training by Dynamic Sparse Model Reconfiguration
- Crossbar-aware neural network pruning
- DFTerNet: Towards 2-bit Dynamic Fusion Networks for Accurate Human Activity Recognition
- Towards Energy-Efficient and Secure Edge AI: A Cross-Layer Framework
- MEC: Memory-efficient Convolution for Deep Neural Network
- Flare: Flexible In-Network Allreduce
- Revisiting Efficient Multi-Step Nonlinearity Compensation with Machine Learning: An Experimental Demonstration
- MobileFaceNets: Efficient CNNs for Accurate Real-Time Face Verification on Mobile Devices
- Efficient Sparse-Winograd Convolutional Neural Networks
- Data Shuffling in Wireless Distributed Computing via Low-Rank Optimization
- BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning Acceleration
- Distilling Object Detectors with Task Adaptive Regularization
- Back to Simplicity: How to Train Accurate BNNs from Scratch?
- Model Pruning Enables Localized and Efficient Federated Learning for Yield Forecasting and Data Sharing
- Compression of Recurrent Neural Networks for Efficient Language Modeling
- DeLTA: GPU Performance Model for Deep Learning Applications with In-depth Memory System Traffic Analysis
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- Delta Networks for Optimized Recurrent Network Computation
- PoPS: Policy Pruning and Shrinking for Deep Reinforcement Learning
- Long short-term memory networks for proton dose calculation in highly heterogeneous tissues
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Multi-Kernel Prediction Networks for Denoising of Burst Images
- Structured Pruning of Deep Convolutional Neural Networks
- ReSpawn: Energy-Efficient Fault-Tolerance for Spiking Neural Networks considering Unreliable Memories
- LiteHAR: Lightweight Human Activity Recognition from WiFi Signals with Random Convolution Kernels
- TinyissimoYOLO: A Quantized, Low-Memory Footprint, TinyML Object Detection Network for Low Power Microcontrollers
- WoodFisher: Efficient Second-Order Approximation for Neural Network Compression
- Complexity-Driven CNN Compression for Resource-constrained Edge AI
- Compact Deep Convolutional Neural Networks With Coarse Pruning
- A Learning Framework for n-bit Quantized Neural Networks toward FPGAs
- Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks
- AxTrain: Hardware-Oriented Neural Network Training for Approximate Inference
- Quality Resilient Deep Neural Networks
- Towards Efficient Training for Neural Network Quantization
- Neural Networks Compression for Language Modeling
- ExPAN(N)D: Exploring Posits for Efficient Artificial Neural Network Design in FPGA-based Systems
- Quantized Convolutional Neural Networks for Mobile Devices
- AnycostFL: Efficient On-Demand Federated Learning over Heterogeneous Edge Devices
- Training Competitive Binary Neural Networks from Scratch
- FPUS23: An Ultrasound Fetus Phantom Dataset with Deep Neural Network Evaluations for Fetus Orientations, Fetal Planes, and Anatomical Features
- Feature Fusion for Online Mutual Knowledge Distillation
- Accurate Optical Flow via Direct Cost Volume Processing
- Learning Sensor Multiplexing Design through Back-propagation
- Towards Transmission-Friendly and Robust CNN Models over Cloud and Device
- Overcoming Challenges in Fixed Point Training of Deep Convolutional Networks
- IGCV3: Interleaved Low-Rank Group Convolutions for Efficient Deep Neural Networks
- Throughput Optimizations for FPGA-based Deep Neural Network Inference
- Learning Sparse Low-Precision Neural Networks With Learnable Regularization
- Efficient Quantized Sparse Matrix Operations on Tensor Cores
- One Size Does Not Fit All: Quantifying and Exposing the Accuracy-Latency Trade-off in Machine Learning Cloud Service APIs via Tolerance Tiers
- MARS: Multi-macro Architecture SRAM CIM-Based Accelerator with Co-designed Compressed Neural Networks
- Task dependent Deep LDA pruning of neural networks
- Run-Time Efficient RNN Compression for Inference on Edge Devices
- Joint Multi-Dimension Pruning via Numerical Gradient Update
- APQ: Joint Search for Network Architecture, Pruning and Quantization Policy
- Probing transfer learning with a model of synthetic correlated datasets
- Towards On-Device AI and Blockchain for 6G enabled Agricultural Supply-chain Management
- Loss-Aware Automatic Selection of Structured Pruning Criteria for Deep Neural Network Acceleration
- Deep Generative Modeling for Scene Synthesis via Hybrid Representations
- BioNetExplorer: Architecture-Space Exploration of Bio-Signal Processing Deep Neural Networks for Wearables
- Degree-Quant: Quantization-Aware Training for Graph Neural Networks
- Touché: Towards Ideal and Efficient Cache Compression By Mitigating Tag Area Overheads
- Improved Techniques for Training Adaptive Deep Networks
- CATBERT: Context-Aware Tiny BERT for Detecting Social Engineering Emails
- Flexible and Fully Quantized Ultra-Lightweight TinyissimoYOLO for Ultra-Low-Power Edge Systems
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Evolutionary Synthesis of Deep Neural Networks via Synaptic Cluster-driven Genetic Encoding
- Optimizing Dense Feed-Forward Neural Networks
- -ARM: Network Sparsification via Stochastic Binary Optimization
- Mixed Precision Low-bit Quantization of Neural Network Language Models for Speech Recognition
- FedGreen: Federated Learning with Fine-Grained Gradient Compression for Green Mobile Edge Computing
- Hardware-Centric AutoML for Mixed-Precision Quantization
- Dynamic Hard Pruning of Neural Networks at the Edge of the Internet
- Transfer Learning for the Efficient Detection of COVID-19 from Smartphone Audio Data
- Few Sample Knowledge Distillation for Efficient Network Compression
- Iteratively Training Look-Up Tables for Network Quantization
- SLAQ: Quality-Driven Scheduling for Distributed Machine Learning
- OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning
- On Quantizing Implicit Neural Representations
- ExSample: Efficient Searches on Video Repositories through Adaptive Sampling
- Biased Mixtures Of Experts: Enabling Computer Vision Inference Under Data Transfer Limitations
- Exploiting Errors for Efficiency: A Survey from Circuits to Algorithms
- KTAN: Knowledge Transfer Adversarial Network
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Channel-wise Mixed-precision Assignment for DNN Inference on Constrained Edge Nodes
- A Novel IoT Trust Model Leveraging Fully Distributed Behavioral Fingerprinting and Secure Delegation
- Trained Rank Pruning for Efficient Deep Neural Networks
- Frank-Wolfe Network: An Interpretable Deep Structure for Non-Sparse Coding
- VEDLIoT: Very Efficient Deep Learning in IoT
- Restricted Recurrent Neural Networks
- A Highly Parallel FPGA Implementation of Sparse Neural Network Training
- Neural network relief: a pruning algorithm based on neural activity
- An Empirical Analysis of the Impact of Data Augmentation on Knowledge Distillation
- Pruning via Iterative Ranking of Sensitivity Statistics
- The Convergence of Machine Learning and Communications
- Exploring Automatic Gym Workouts Recognition Locally On Wearable Resource-Constrained Devices
- LCNN: Lookup-based Convolutional Neural Network
- BoolNet: Minimizing The Energy Consumption of Binary Neural Networks
- Accelerating and Compressing Deep Neural Networks for Massive MIMO CSI Feedback
- FedDCT: Federated Learning of Large Convolutional Neural Networks on Resource Constrained Devices using Divide and Collaborative Training
- Sauron U-Net: Simple automated redundancy elimination in medical image segmentation via filter pruning
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
- Intermittent Learning: On-Device Machine Learning on Intermittently Powered System
- JSDoop and TensorFlow.js: Volunteer Distributed Web Browser-Based Neural Network Training
- Automated Design Space Exploration for optimised Deployment of DNN on Arm Cortex-A CPUs
- ECM: Early Exit via Class Means for Efficient Supervised and Unsupervised Learning
- Constrained Deep Learning using Conditional Gradient and Applications in Computer Vision
- Coordinating Filters for Faster Deep Neural Networks
- FxP-QNet: A Post-Training Quantizer for the Design of Mixed Low-Precision DNNs with Dynamic Fixed-Point Representation
- Understanding Generalization in Deep Learning via Tensor Methods
- Progressive Neural Compression for Adaptive Image Offloading under Timing Constraints
- On-FPGA Training with Ultra Memory Reduction: A Low-Precision Tensor Method
- Tensor Contraction Layers for Parsimonious Deep Nets
- Edge-Cloud Cooperation for DNN Inference via Reinforcement Learning and Supervised Learning
- Real-Time Execution of Large-scale Language Models on Mobile
- SymbolNet: Neural Symbolic Regression with Adaptive Dynamic Pruning for Compression
- Applying the Residue Number System to Network Inference
- Federated Learning for Commercial Image Sources
- Improving Neural Network Training using Dynamic Learning Rate Schedule for PINNs and Image Classification
- ASP Vision: Optically Computing the First Layer of Convolutional Neural Networks using Angle Sensitive Pixels
- Joint Pruning & Quantization for Extremely Sparse Neural Networks
- Accelerating Monte Carlo Bayesian Inference via Approximating Predictive Uncertainty over Simplex
- Depthwise Multiception Convolution for Reducing Network Parameters without Sacrificing Accuracy
- HEMP: High-order Entropy Minimization for neural network comPression
- Channel-wise Hessian Aware trace-Weighted Quantization of Neural Networks
- Improving Neural Network with Uniform Sparse Connectivity
- SmartSplit: Latency-Energy-Memory Optimisation for CNN Splitting on Smartphone Environment
- Espresso: Efficient Forward Propagation for BCNNs
- Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
- Memristor-based Deep Convolution Neural Network: A Case Study
- Topological Persistence Guided Knowledge Distillation for Wearable Sensor Data
- PTEENet: Post-Trained Early-Exit Neural Networks Augmentation for Inference Cost Optimization
- Latency-Constrained DNN Architecture Learning for Edge Systems using Zerorized Batch Normalization
- Learning Versatile Convolution Filters for Efficient Visual Recognition
- Low-Power Object Counting with Hierarchical Neural Networks
- PURSUhInT: In Search of Informative Hint Points Based on Layer Clustering for Knowledge Distillation
- Interleaver Design for Deep Neural Networks
- Streamlining Tensor and Network Pruning in PyTorch
- Deep Spiking Neural Networks for Large Vocabulary Automatic Speech Recognition
- PCNN: Pattern-based Fine-Grained Regular Pruning towards Optimizing CNN Accelerators
- At-Scale Evaluation of Weight Clustering to Enable Energy-Efficient Object Detection
- Improving Neural Network Training in Low Dimensional Random Bases
- Harmonic Networks: Integrating Spectral Information into CNNs
- On the efficient representation and execution of deep acoustic models
- BiQGEMM: Matrix Multiplication with Lookup Table For Binary-Coding-based Quantized DNNs
- Enabling Embedded Inference Engine with ARM Compute Library: A Case Study
- Demystifying Neural Network Filter Pruning
- Weight Pruning via Adaptive Sparsity Loss
- Partition Pruning: Parallelization-Aware Pruning for Deep Neural Networks
- Convergence of a Relaxed Variable Splitting Method for Learning Sparse Neural Networks via , and transformed- Penalties
- Compressing Gradient Optimizers via Count-Sketches
- Pufferfish: Communication-efficient Models At No Extra Cost
- LegoNet: Memory Footprint Reduction Through Block Weight Clustering
- Hardware-software co-exploration with racetrack memory based in-memory computing for CNN inference in embedded systems
- HG-Caffe: Mobile and Embedded Neural Network GPU (OpenCL) Inference Engine with FP16 Supporting
- Learning Efficient Convolutional Networks through Irregular Convolutional Kernels
- Knowledge Distillation Under Ideal Joint Classifier Assumption
- Exploiting Channel Similarity for Accelerating Deep Convolutional Neural Networks
- SparCE: Sparsity aware General Purpose Core Extensions to Accelerate Deep Neural Networks
- Compressed Learning of Deep Neural Networks for OpenCL-Capable Embedded Systems
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- Exploring Weight Importance and Hessian Bias in Model Pruning
- Powers of layers for image-to-image translation
- High Performance Convolution Using Sparsity and Patterns for Inference in Deep Convolutional Neural Networks
- CRAFT: Criticality-Aware Fault-Tolerance Enhancement Techniques for Emerging Memories-Based Deep Neural Networks
- Statistical Model Compression for Small-Footprint Natural Language Understanding
- QuickNet: Maximizing Efficiency and Efficacy in Deep Architectures
- STH: Spatio-Temporal Hybrid Convolution for Efficient Action Recognition
- A New Clustering-Based Technique for the Acceleration of Deep Convolutional Networks
- GroupReduce: Block-Wise Low-Rank Approximation for Neural Language Model Shrinking
- Graph DNA: Deep Neighborhood Aware Graph Encoding for Collaborative Filtering
- Jointly Optimizing Preprocessing and Inference for DNN-based Visual Analytics
- Pruning and Slicing Neural Networks using Formal Verification
- Distilling with Performance Enhanced Students
- A Winning Hand: Compressing Deep Networks Can Improve Out-Of-Distribution Robustness
- On Iterative Neural Network Pruning, Reinitialization, and the Similarity of Masks
- REPrune: Filter Pruning via Representative Election
- SEVEN: Pruning Transformer Model by Reserving Sentinels
- Pruning Convolutional Neural Networks for Image Instance Retrieval
- HALP: Hardware-Aware Latency Pruning
- Efficient Integer-Arithmetic-Only Convolutional Neural Networks
- Visual Confusion Label Tree For Image Classification
- Graph Neural Networks Including Sparse Interpretability
- Active Subspace of Neural Networks: Structural Analysis and Universal Attacks
- Neural Machine Translation with 4-Bit Precision and Beyond
- Learning Sparse & Ternary Neural Networks with Entropy-Constrained Trained Ternarization (EC2T)
- Model order reduction with neural networks: Application to laminar and turbulent flows
- Activation Density driven Energy-Efficient Pruning in Training
- Multi-modal Experts Network for Autonomous Driving
- Hierarchical compositional feature learning
- Automated Backend-Aware Post-Training Quantization
- Network Automatic Pruning: Start NAP and Take a Nap
- EasyConvPooling: Random Pooling with Easy Convolution for Accelerating Training and Testing
- DCentNet: Decentralized Multistage Biomedical Signal Classification using Early Exits
- RTN: Reparameterized Ternary Network
- Leveraging Structured Pruning of Convolutional Neural Networks
- RISC-NN: Use RISC, NOT CISC as Neural Network Hardware Infrastructure
- Extending Label Smoothing Regularization with Self-Knowledge Distillation
- Compression and Localization in Reinforcement Learning for ATARI Games
- Efficient Inferencing of Compressed Deep Neural Networks
- ATCN: Resource-Efficient Processing of Time Series on Edge
- Bespoke vs. Prêt-à-Porter Lottery Tickets: Exploiting Mask Similarity for Trainable Sub-Network Finding
- Fixed-point optimization of deep neural networks with adaptive step size retraining
- Training Sparse Neural Networks using Compressed Sensing
- Exploiting Weight Redundancy in CNNs: Beyond Pruning and Quantization
- Towards Collaborative Intelligence Friendly Architectures for Deep Learning
- Latency-Memory Optimized Splitting of Convolution Neural Networks for Resource Constrained Edge Devices
- Blind Adversarial Pruning: Balance Accuracy, Efficiency and Robustness
- Rethinking Empirical Evaluation of Adversarial Robustness Using First-Order Attack Methods
- Towards Accurate Quantization and Pruning via Data-free Knowledge Transfer
- An Approximation Algorithm for Optimal Subarchitecture Extraction
- Data-free mixed-precision quantization using novel sensitivity metric
- Mitigate Parasitic Resistance in Resistive Crossbar-based Convolutional Neural Networks
- AutoPruning for Deep Neural Network with Dynamic Channel Masking
- Towards Accurate and High-Speed Spiking Neuromorphic Systems with Data Quantization-Aware Deep Networks
- Sparse Persistent RNNs: Squeezing Large Recurrent Networks On-Chip
- Mining the Weights Knowledge for Optimizing Neural Network Structures
- A Survey on GAN Acceleration Using Memory Compression Technique
- MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
- MOFHEI: Model Optimizing Framework for Fast and Efficient Homomorphically Encrypted Neural Network Inference
- FLASH: Fast Neural Architecture Search with Hardware Optimization
- A scalable convolutional neural network for task-specified scenarios via knowledge distillation
- Accelerator-Aware Training for Transducer-Based Speech Recognition
- SIPA: A Simple Framework for Efficient Networks
- Filter Pre-Pruning for Improved Fine-tuning of Quantized Deep Neural Networks
- Multi-objective Evolutionary Approach for Efficient Kernel Size and Shape for CNN
- KCP: Kernel Cluster Pruning for Dense Labeling Neural Networks
- A Comparative Study of Neural Network Compression
- The Future of Intelligent Wavefront Shaping for Smart Radio Environments
- SmartDeal: Re-Modeling Deep Network Weights for Efficient Inference and Training
- Deep Compression for PyTorch Model Deployment on Microcontrollers
- Grow-Push-Prune: aligning deep discriminants for effective structural network compression
- Stacked Filters Stationary Flow For Hardware-Oriented Acceleration Of Deep Convolutional Neural Networks
- Prune the Convolutional Neural Networks with Sparse Shrink
- clcNet: Improving the Efficiency of Convolutional Neural Network using Channel Local Convolutions
- Class-Discriminative CNN Compression
- Supporting Massive DLRM Inference Through Software Defined Memory
- FPGA Implementations of 3D-SIMD Processor Architecture for Deep Neural Networks Using Relative Indexed Compressed Sparse Filter Encoding Format and Stacked Filters Stationary Flow
- Constrained-size Tensorflow Models for YouTube-8M Video Understanding Challenge