DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
arXiv:1606.06160
Abstract
We propose DoReFa-Net, a method to train convolutional neural networks that have low bitwidth weights and activations using low bitwidth parameter gradients. In particular, during backward pass, parameter gradients are stochastically quantized to low bitwidth numbers before being propagated to convolutional layers. As convolutions during forward/backward passes can now operate on low bitwidth weights and activations/gradients respectively, DoReFa-Net can use bit convolution kernels to accelerate both training and inference. Moreover, as bit convolutions can be efficiently implemented on CPU, FPGA, ASIC and GPU, DoReFa-Net opens the way to accelerate training of low bitwidth neural network on these hardware. Our experiments on SVHN and ImageNet datasets prove that DoReFa-Net can achieve comparable prediction accuracy as 32-bit counterparts. For example, a DoReFa-Net derived from AlexNet that has 1-bit weights, 2-bit activations, can be trained from scratch using 6-bit gradients to get 46.1\% top-1 accuracy on ImageNet validation set. The DoReFa-Net AlexNet model is released publicly.
References in corpus (5)
- Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
- Deep Learning with Limited Numerical Precision
- Compressing Deep Convolutional Networks using Vector Quantization
- Ternary Weight Networks
- Deep neural networks are robust to weight binarization and other non-linear distortions
Cited by in corpus (477)
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
- Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- Trained Ternary Quantization
- PACT: Parameterized Clipping Activation for Quantized Neural Networks
- Machine Learning at the Wireless Edge: Distributed Stochastic Gradient Descent Over-the-Air
- Binary Neural Networks: A Survey
- FPGA-based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- Training and Inference with Integers in Deep Neural Networks
- NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
- Learned Step Size Quantization
- MCUNet: Tiny Deep Learning on IoT Devices
- Single Path One-Shot Neural Architecture Search with Uniform Sampling
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead
- Don't Use Large Mini-Batches, Use Local SGD
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Gradient Sparsification for Communication-Efficient Distributed Optimization
- A Survey on Methods and Theories of Quantized Neural Networks
- ATOMO: Communication-efficient Learning via Atomic Sparsification
- Scalable Methods for 8-bit Training of Neural Networks
- Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors
- Mixed Precision Quantization of ConvNets via Differentiable Neural Architecture Search
- WRPN: Wide Reduced-Precision Networks
- Tiny Machine Learning: Progress and Futures
- Bringing AI To Edge: From Deep Learning's Perspective
- Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks
- Discrimination-aware Channel Pruning for Deep Neural Networks
- Discrimination-aware Network Pruning for Deep Model Compression
- OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
- A Survey of FPGA-Based Neural Network Accelerator
- Hardware Approximate Techniques for Deep Neural Network Accelerators: A Survey
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- XNOR Neural Engine: a Hardware Accelerator IP for 21.6 fJ/op Binary Neural Network Inference
- Post-training 4-bit quantization of convolution networks for rapid-deployment
- Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural Networks
- XNOR-Net++: Improved Binary Neural Networks
- Scaling Distributed Machine Learning with In-Network Aggregation
- Exploring the Connection Between Binary and Spiking Neural Networks
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights
- GraphVite: A High-Performance CPU-GPU Hybrid System for Node Embedding
- Layer-specific Optimization for Mixed Data Flow with Mixed Precision in FPGA Design for CNN-based Object Detectors
- Training Quantized Nets: A Deeper Understanding
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- Loss-aware Weight Quantization of Deep Networks
- Energy-efficient stochastic computing with superparamagnetic tunnel junctions
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- Deep Convolutional Neural Network Inference with Floating-point Weights and Fixed-point Activations
- BS4NN: Binarized Spiking Neural Networks with Temporal Coding and Learning
- Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
- Efficient and Robust Parallel DNN Training through Model Parallelism on Multi-GPU Platform
- Applications and Techniques for Fast Machine Learning in Science
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- MeliusNet: Can Binary Neural Networks Achieve MobileNet-level Accuracy?
- Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming
- ProxQuant: Quantized Neural Networks via Proximal Operators
- Training Binary Neural Networks with Real-to-Binary Convolutions
- AddNet: Deep Neural Networks Using FPGA-Optimized Multipliers
- Latent Weights Do Not Exist: Rethinking Binarized Neural Network Optimization
- Rotated Binary Neural Network
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- And the Bit Goes Down: Revisiting the Quantization of Neural Networks
- SLSNet: Skin lesion segmentation using a lightweight generative adversarial network
- Effective Quantization Methods for Recurrent Neural Networks
- Ternary Neural Networks with Fine-Grained Quantization
- Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks
- Trained Quantization Thresholds for Accurate and Efficient Fixed-Point Inference of Deep Neural Networks
- Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
- Charged particle tracking via edge-classifying interaction networks
- Neural Network Distiller: A Python Package For DNN Compression Research
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
- Guided Hybrid Quantization for Object detection in Multimodal Remote Sensing Imagery via One-to-one Self-teaching
- MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning
- DFTerNet: Towards 2-bit Dynamic Fusion Networks for Accurate Human Activity Recognition
- HAQ: Hardware-Aware Automated Quantization with Mixed Precision
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- Constructing Energy-efficient Mixed-precision Neural Networks through Principal Component Analysis for Edge Intelligence
- Efficient Processing of Deep Neural Networks: A Tutorial and Survey
- Model Rubik's Cube: Twisting Resolution, Depth and Width for TinyNets
- QKD: Quantization-aware Knowledge Distillation
- BinaryBERT: Pushing the Limit of BERT Quantization
- Back to Simplicity: How to Train Accurate BNNs from Scratch?
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- ReLeQ: A Reinforcement Learning Approach for Deep Quantization of Neural Networks
- A Unified Framework of DNN Weight Pruning and Weight Clustering/Quantization Using ADMM
- Mixed Precision Training With 8-bit Floating Point
- A Survey of FPGA Based Deep Learning Accelerators: Challenges and Opportunities
- Improved training of binary networks for human pose estimation and image recognition
- Hierarchical binary CNNs for landmark localization with limited resources
- Epilepsy Seizure Detection and Prediction using an Approximate Spiking Convolutional Transformer
- NITI: Training Integer Neural Networks Using Integer-only Arithmetic
- More is Less: A More Complicated Network with Less Inference Complexity
- Quantization of Fully Convolutional Networks for Accurate Biomedical Image Segmentation
- Bridging the Accuracy Gap for 2-bit Quantized Neural Networks (QNN)
- HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision
- NICE: Noise Injection and Clamping Estimation for Neural Network Quantization
- Defensive Quantization: When Efficiency Meets Robustness
- Robust Quantization: One Model to Rule Them All
- Searching for Low-Bit Weights in Quantized Neural Networks
- A Learning Framework for n-bit Quantized Neural Networks toward FPGAs
- Training Binary Neural Networks through Learning with Noisy Supervision
- Defend Deep Neural Networks Against Adversarial Examples via Fixed and Dynamic Quantized Activation Functions
- Towards Efficient Training for Neural Network Quantization
- ExPAN(N)D: Exploring Posits for Efficient Artificial Neural Network Design in FPGA-based Systems
- Accurate and Compact Convolutional Neural Networks with Trained Binarization
- Controlling Information Capacity of Binary Neural Network
- DSConv: Efficient Convolution Operator
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary Activations
- Training Competitive Binary Neural Networks from Scratch
- FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural Networks
- Bit Fusion: Bit-Level Dynamically Composable Architecture for Accelerating Deep Neural Networks
- Adaptive Gradient Quantization for Data-Parallel SGD
- ReActNet: Towards Precise Binary Neural Network with Generalized Activation Functions
- Energy Efficient Learning with Low Resolution Stochastic Domain Wall Synapse Based Deep Neural Networks
- Efficient Exact Verification of Binarized Neural Networks
- Streamlined Deployment for Quantized Neural Networks
- Binary Neural Networks for Memory-Efficient and Effective Visual Place Recognition in Changing Environments
- IGCV3: Interleaved Low-Rank Group Convolutions for Efficient Deep Neural Networks
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
- Kernel Based Progressive Distillation for Adder Neural Networks
- Learning Sparse Low-Precision Neural Networks With Learnable Regularization
- Model compression as constrained optimization, with application to neural nets. Part I: general framework
- Learning Accurate Low-Bit Deep Neural Networks with Stochastic Quantization
- Training DNNs with Hybrid Block Floating Point
- Forward and Backward Information Retention for Accurate Binary Neural Networks
- Neural Network-Optimized Channel Estimator and Training Signal Design for MIMO Systems with Few-Bit ADCs
- MARS: Multi-macro Architecture SRAM CIM-Based Accelerator with Co-designed Compressed Neural Networks
- Bi-Real Net: Enhancing the Performance of 1-bit CNNs With Improved Representational Capability and Advanced Training Algorithm
- Self-Binarizing Networks
- Model compression as constrained optimization, with application to neural nets. Part II: quantization
- Mixed Low-precision Deep Learning Inference using Dynamic Fixed Point
- Attacking Binarized Neural Networks
- Towards Effective Low-bitwidth Convolutional Neural Networks
- AdaptivFloat: A Floating-point based Data Type for Resilient Deep Learning Inference
- QGAN: Quantized Generative Adversarial Networks
- DeepTwist: Learning Model Compression via Occasional Weight Distortion
- GhostSR: Learning Ghost Features for Efficient Image Super-Resolution
- Quantune: Post-training Quantization of Convolutional Neural Networks using Extreme Gradient Boosting for Fast Deployment
- NAND-SPIN-Based Processing-in-MRAM Architecture for Convolutional Neural Network Acceleration
- Resource-Efficient Neural Networks for Embedded Systems
- Regularizing Activation Distribution for Training Binarized Deep Networks
- Fixed-point Quantization of Convolutional Neural Networks for Quantized Inference on Embedded Platforms
- Learning Convolutional Networks for Content-weighted Image Compression
- LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression
- Deep Learning for Semantic Segmentation on Minimal Hardware
- ZeroQ: A Novel Zero Shot Quantization Framework
- Role-Wise Data Augmentation for Knowledge Distillation
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Projection Convolutional Neural Networks for 1-bit CNNs via Discrete Back Propagation
- DeepHammer: Depleting the Intelligence of Deep Neural Networks through Targeted Chain of Bit Flips
- HAWQV3: Dyadic Neural Network Quantization
- Deep Learning as a Mixed Convex-Combinatorial Optimization Problem
- N3H-Core: Neuron-designed Neural Network Accelerator via FPGA-based Heterogeneous Computing Cores
- Robust Sparse Regularization: Simultaneously Optimizing Neural Network Robustness and Compactness
- Compressive Sensing Using Iterative Hard Thresholding with Low Precision Data Representation: Theory and Applications
- AdaBits: Neural Network Quantization with Adaptive Bit-Widths
- Quantization of Deep Neural Networks for Accumulator-constrained Processors
- Towards the AlexNet Moment for Homomorphic Encryption: HCNN, theFirst Homomorphic CNN on Encrypted Data with GPUs
- Learning to Train a Binary Neural Network
- Iteratively Training Look-Up Tables for Network Quantization
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- Relaxed Quantization for Discretized Neural Networks
- I-BERT: Integer-only BERT Quantization
- Post-Training 4-bit Quantization on Embedding Tables
- Bit-Flip Attack: Crushing Neural Network with Progressive Bit Search
- Hardware-Centric AutoML for Mixed-Precision Quantization
- Switchable Precision Neural Networks
- Post-Training Piecewise Linear Quantization for Deep Neural Networks
- Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
- Sparse Communication for Training Deep Networks
- MTP: Multi-Task Pruning for Efficient Semantic Segmentation Networks
- MoBiNet: A Mobile Binary Network for Image Classification
- The ZipML Framework for Training Models with End-to-End Low Precision: The Cans, the Cannots, and a Little Bit of Deep Learning
- Hierarchical Training of Deep Neural Networks Using Early Exiting
- A Statistical Framework for Low-bitwidth Training of Deep Neural Networks
- AdderNet: Do We Really Need Multiplications in Deep Learning?
- Generative Low-bitwidth Data Free Quantization
- A Data and Compute Efficient Design for Limited-Resources Deep Learning
- DarKnight: A Data Privacy Scheme for Training and Inference of Deep Neural Networks
- BatchQuant: Quantized-for-all Architecture Search with Robust Quantizer
- Mirage: An RNS-Based Photonic Accelerator for DNN Training
- VS-Quant: Per-vector Scaled Quantization for Accurate Low-Precision Neural Network Inference
- Efficient Bitwidth Search for Practical Mixed Precision Neural Network
- Ternary Compression for Communication-Efficient Federated Learning
- QPyTorch: A Low-Precision Arithmetic Simulation Framework
- SpikeGrad: An ANN-equivalent Computation Model for Implementing Backpropagation with Spikes
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- Comprehensive SNN Compression Using ADMM Optimization and Activity Regularization
- DaCapo: Accelerating Continuous Learning in Autonomous Systems for Video Analytics
- EMPIR: Ensembles of Mixed Precision Deep Networks for Increased Robustness against Adversarial Attacks
- Multi-Prize Lottery Ticket Hypothesis: Finding Accurate Binary Neural Networks by Pruning A Randomly Weighted Network
- TBT: Targeted Neural Network Attack with Bit Trojan
- Blended Coarse Gradient Descent for Full Quantization of Deep Neural Networks
- BoolNet: Minimizing The Energy Consumption of Binary Neural Networks
- Quantized Densely Connected U-Nets for Efficient Landmark Localization
- Toolflows for Mapping Convolutional Neural Networks on FPGAs: A Survey and Future Directions
- Fix your classifier: the marginal value of training the last weight layer
- Retraining-Based Iterative Weight Quantization for Deep Neural Networks
- Rethinking Differentiable Search for Mixed-Precision Neural Networks
- PAMS: Quantized Super-Resolution via Parameterized Max Scale
- Bit Efficient Quantization for Deep Neural Networks
- FxP-QNet: A Post-Training Quantizer for the Design of Mixed Low-Precision DNNs with Dynamic Fixed-Point Representation
- An Information Theory-inspired Strategy for Automatic Network Pruning
- BiFSMN: Binary Neural Network for Keyword Spotting
- GeneCAI: Genetic Evolution for Acquiring Compact AI
- CPT: Efficient Deep Neural Network Training via Cyclic Precision
- BMXNet: An Open-Source Binary Neural Network Implementation Based on MXNet
- ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training
- ApproxNet: Content and Contention-Aware Video Analytics System for Embedded Clients
- BiPointNet: Binary Neural Network for Point Clouds
- Mirror Descent View for Neural Network Quantization
- Sparse Weight Activation Training
- GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training
- Precision Highway for Ultra Low-Precision Quantization
- Rethinking Class-Discrimination Based CNN Channel Pruning
- Weight Normalization based Quantization for Deep Neural Network Compression
- Applying the Residue Number System to Network Inference
- Training High-Performance and Large-Scale Deep Neural Networks with Full 8-bit Integers
- Full deep neural network training on a pruned weight budget
- Efficient Deep Neural Networks
- CAT: Compression-Aware Training for bandwidth reduction
- Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
- Channel-wise Hessian Aware trace-Weighted Quantization of Neural Networks
- Per-Tensor Fixed-Point Quantization of the Back-Propagation Algorithm
- Training Compact Neural Networks with Binary Weights and Low Precision Activations
- SQuantizer: Simultaneous Learning for Both Sparse and Low-precision Neural Networks
- A Generalized Zero-Shot Quantization of Deep Convolutional Neural Networks via Learned Weights Statistics
- Learning Frequency Domain Approximation for Binary Neural Networks
- BinaryRelax: A Relaxation Approach For Training Deep Neural Networks With Quantized Weights
- HEMP: High-order Entropy Minimization for neural network comPression
- Scheduling Policy and Power Allocation for Federated Learning in NOMA Based MEC
- Any-Precision Deep Neural Networks
- SmartExchange: Trading Higher-cost Memory Storage/Access for Lower-cost Computation
- Training Bit Fully Convolutional Network for Fast Semantic Segmentation
- CodeX: Bit-Flexible Encoding for Streaming-based FPGA Acceleration of DNNs
- Network Quantization with Element-wise Gradient Scaling
- Boosting Binary Masks for Multi-Domain Learning through Affine Transformations
- Improving Adversarial Robustness in Weight-quantized Neural Networks
- Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
- Sharpness-aware Quantization for Deep Neural Networks
- Elastic Significant Bit Quantization and Acceleration for Deep Neural Networks
- Revisiting BFloat16 Training
- Bidirectional compression in heterogeneous settings for distributed or federated learning with partial participation: tight convergence guarantees
- Bit Error Robustness for Energy-Efficient DNN Accelerators
- Accelerating CNN Training by Pruning Activation Gradients
- Exploiting Kernel Sparsity and Entropy for Interpretable CNN Compression
- Bi-Real Net: Binarizing Deep Network Towards Real-Network Performance
- Simultaneously Optimizing Weight and Quantizer of Ternary Neural Network using Truncated Gaussian Approximation
- WRPN: Training and Inference using Wide Reduced-Precision Networks
- Convolutional Neural Networks Quantization with Attention
- Deep Spiking Neural Networks for Large Vocabulary Automatic Speech Recognition
- MWQ: Multiscale Wavelet Quantized Neural Networks
- Effective and Fast: A Novel Sequential Single Path Search for Mixed-Precision Quantization
- Automating Generation of Low Precision Deep Learning Operators
- Rethinking Floating Point Overheads for Mixed Precision DNN Accelerators
- Fractional Skipping: Towards Finer-Grained Dynamic CNN Inference
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
- Robustness and Transferability of Universal Attacks on Compressed Models
- Distillation Guided Residual Learning for Binary Convolutional Neural Networks
- Towards Lossless Binary Convolutional Neural Networks Using Piecewise Approximation
- Hu-Fu: Hardware and Software Collaborative Attack Framework against Neural Networks
- Propagating Asymptotic-Estimated Gradients for Low Bitwidth Quantized Neural Networks
- Collaborative Distillation for Ultra-Resolution Universal Style Transfer
- Accelerated Training for CNN Distributed Deep Learning through Automatic Resource-Aware Layer Placement
- Efficient Halftoning via Deep Reinforcement Learning
- TaxoNN: A Light-Weight Accelerator for Deep Neural Network Training
- Condensation-Net: Memory-Efficient Network Architecture with Cross-Channel Pooling Layers and Virtual Feature Maps
- Scaling Binarized Neural Networks on Reconfigurable Logic
- Taming Binarized Neural Networks and Mixed-Integer Programs
- WrapNet: Neural Net Inference with Ultra-Low-Resolution Arithmetic
- Post-training Quantization with Multiple Points: Mixed Precision without Mixed Precision
- Ternary Residual Networks
- Evolutionary Bin Packing for Memory-Efficient Dataflow Inference Acceleration on FPGA
- Neural Network Compression Via Sparse Optimization
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation
- Divide and Conquer: Leveraging Intermediate Feature Representations for Quantized Training of Neural Networks
- Differentiable Dynamic Quantization with Mixed Precision and Adaptive Resolution
- Conditionally Deep Hybrid Neural Networks Across Edge and Cloud
- Ordering Chaos: Memory-Aware Scheduling of Irregularly Wired Neural Networks for Edge Devices
- Binarizing by Classification: Is soft function really necessary?
- Learning Content-Weighted Deep Image Compression
- Design Flow of Accelerating Hybrid Extremely Low Bit-width Neural Network in Embedded FPGA
- Adaptive Precision Training: Quantify Back Propagation in Neural Networks with Fixed-point Numbers
- Direction is what you need: Improving Word Embedding Compression in Large Language Models
- Model compression as constrained optimization, with application to neural nets. Part V: combining compressions
- BitNet: Bit-Regularized Deep Neural Networks
- Learning Architectures for Binary Networks
- Qimera: Data-free Quantization with Synthetic Boundary Supporting Samples
- Improving Accuracy of Binary Neural Networks using Unbalanced Activation Distribution
- An efficient and perceptually motivated auditory neural encoding and decoding algorithm for spiking neural networks
- DEEPEYE: A Compact and Accurate Video Comprehension at Terminal Devices Compressed with Quantization and Tensorization
- Generative Zero-shot Network Quantization
- QuTiBench: Benchmarking Neural Networks on Heterogeneous Hardware
- Mixed-Precision Quantized Neural Network with Progressively Decreasing Bitwidth For Image Classification and Object Detection
- NUQSGD: Provably Communication-efficient Data-parallel SGD via Nonuniform Quantization
- FIXAR: A Fixed-Point Deep Reinforcement Learning Platform with Quantization-Aware Training and Adaptive Parallelism
- Pufferfish: Communication-efficient Models At No Extra Cost
- SinReQ: Generalized Sinusoidal Regularization for Low-Bitwidth Deep Quantized Training
- Towards Learning of Filter-Level Heterogeneous Compression of Convolutional Neural Networks
- Learning Context-Based Non-local Entropy Modeling for Image Compression
- Dynamic Runtime Feature Map Pruning
- A flexible, extensible software framework for model compression based on the LC algorithm
- Binarizing MobileNet via Evolution-based Searching
- Dataflow-based Joint Quantization of Weights and Activations for Deep Neural Networks
- Low-Rank Training of Deep Neural Networks for Emerging Memory Technology
- PIMBALL: Binary Neural Networks in Spintronic Memory
- Quantization of Deep Neural Networks for Accurate Edge Computing
- No Multiplication? No Floating Point? No Problem! Training Networks for Efficient Inference
- Training Deep Neural Network in Limited Precision
- Knowledge distillation for optimization of quantized deep neural networks
- FAT: Learning Low-Bitwidth Parametric Representation via Frequency-Aware Transformation
- Predicting Adversarial Examples with High Confidence
- Binarized Neural Architecture Search for Efficient Object Recognition
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network Quantization
- Intermittent Pulling with Local Compensation for Communication-Efficient Federated Learning
- Dynamic Sparse Graph for Efficient Deep Learning
- MSP: An FPGA-Specific Mixed-Scheme, Multi-Precision Deep Neural Network Quantization Framework
- Matrix and tensor decompositions for training binary neural networks
- A Winning Hand: Compressing Deep Networks Can Improve Out-Of-Distribution Robustness
- Efficient non-uniform quantizer for quantized neural network targeting reconfigurable hardware
- Auto-tuning Neural Network Quantization Framework for Collaborative Inference Between the Cloud and Edge
- MetaMix: Meta-state Precision Searcher for Mixed-precision Activation Quantization
- A High-Performance Adaptive Quantization Approach for Edge CNN Applications
- Optimal Gradient Quantization Condition for Communication-Efficient Distributed Training
- Alternating Direction Method of Multipliers for Quantization
- A Real-time Low-cost Artificial Intelligence System for Autonomous Spraying in Palm Plantations
- Learning Recurrent Binary/Ternary Weights
- Recent Advances in Efficient Computation of Deep Convolutional Neural Networks
- Quantized Convolutional Neural Networks Through the Lens of Partial Differential Equations
- An FPGA Accelerated Method for Training Feed-forward Neural Networks Using Alternating Direction Method of Multipliers and LSMR
- Dynamic Network Quantization for Efficient Video Inference
- One Model for All Quantization: A Quantized Network Supporting Hot-Swap Bit-Width Adjustment
- n-hot: Efficient bit-level sparsity for powers-of-two neural network quantization
- RA-BNN: Constructing Robust & Accurate Binary Neural Network to Simultaneously Defend Adversarial Bit-Flip Attack and Improve Accuracy
- Optimal Quantization for Batch Normalization in Neural Network Deployments and Beyond
- Improving Binary Neural Networks through Fully Utilizing Latent Weights
- Merging and Evolution: Improving Convolutional Neural Networks for Mobile Applications
- High-Capacity Expert Binary Networks
- NASB: Neural Architecture Search for Binary Convolutional Neural Networks
- Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural Networks
- Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural Networks
- Search What You Want: Barrier Panelty NAS for Mixed Precision Quantization
- Towards Fast and Energy-Efficient Binarized Neural Network Inference on FPGA
- Distributed Low Precision Training Without Mixed Precision
- Sampling-Free Learning of Bayesian Quantized Neural Networks
- Progressive Learning of Low-Precision Networks
- Table-Based Neural Units: Fully Quantizing Networks for Multiply-Free Inference
- An Investigation on Different Underlying Quantization Schemes for Pre-trained Language Models
- Distribution-sensitive Information Retention for Accurate Binary Neural Network
- Histogram-Equalized Quantization for logic-gated Residual Neural Networks
- Continual Learning via Bit-Level Information Preserving
- Path Sample-Analytic Gradient Estimators for Stochastic Binary Networks
- Efficient Integer-Arithmetic-Only Convolutional Neural Networks
- RTN: Reparameterized Ternary Network
- Processing-In-Memory Acceleration of Convolutional Neural Networks for Energy-Efficiency, and Power-Intermittency Resilience
- Quantized Adam with Error Feedback
- DNN Feature Map Compression using Learned Representation over GF(2)
- BCNN: Binary Complex Neural Network
- LiMuSE: Lightweight Multi-modal Speaker Extraction
- Dense xUnit Networks
- Training Quantized Neural Networks with a Full-precision Auxiliary Module
- Generative Design of Hardware-aware DNNs
- Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions
- Fast Adjustable Threshold For Uniform Neural Network Quantization (Winning solution of LPIRC-II)
- Hardware-friendly Neural Network Architecture for Neuromorphic Computing
- A Layer-wise Adversarial-aware Quantization Optimization for Improving Robustness
- BNAS v2: Learning Architectures for Binary Networks with Empirical Improvements
- SiMaN: Sign-to-Magnitude Network Binarization
- Running Neural Networks on the NIC
- AQD: Towards Accurate Fully-Quantized Object Detection
- Universal Adder Neural Networks
- How Low Can We Go: Trading Memory for Error in Low-Precision Training
- RRAM based neuromorphic algorithms
- Efficient Crowd Counting via Structured Knowledge Transfer
- ReCU: Reviving the Dead Weights in Binary Neural Networks
- Trajectory Normalized Gradients for Distributed Optimization
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- Balanced Binary Neural Networks with Gated Residual
- Recent Advances in Convolutional Neural Network Acceleration
- Post-Training Sparsity-Aware Quantization
- Auto-NBA: Efficient and Effective Search Over the Joint Space of Networks, Bitwidths, and Accelerators
- A Very Compact Embedded CNN Processor Design Based on Logarithmic Computing
- WaveQ: Gradient-Based Deep Quantization of Neural Networks through Sinusoidal Adaptive Regularization
- Searching for Accurate Binary Neural Architectures
- GDRQ: Group-based Distribution Reshaping for Quantization
- Composite Binary Decomposition Networks
- Improved Techniques for Quantizing Deep Networks with Adaptive Bit-Widths
- Enabling Binary Neural Network Training on the Edge
- Pruning Ternary Quantization
- Rapid Elastic Architecture Search under Specialized Classes and Resource Constraints
- Arch-Net: Model Distillation for Architecture Agnostic Model Deployment
- HERO: Hessian-Enhanced Robust Optimization for Unifying and Improving Generalization and Quantization Performance
- Training of mixed-signal optical convolutional neural network with reduced quantization level
- FATNN: Fast and Accurate Ternary Neural Networks
- Hessian-aware Quantized Node Embeddings for Recommendation
- Pyramid Vector Quantization for Deep Learning
- CoDeNet: Efficient Deployment of Input-Adaptive Object Detection on Embedded FPGAs
- Is In-Domain Data Really Needed? A Pilot Study on Cross-Domain Calibration for Network Quantization
- Efficient Micro-Structured Weight Unification and Pruning for Neural Network Compression
- BiDet: An Efficient Binarized Object Detector
- Complexity-aware Adaptive Training and Inference for Edge-Cloud Distributed AI Systems
- Hardware and software co-optimization for the initialization failure of the ReRAM based cross-bar array
- ILMPQ : An Intra-Layer Multi-Precision Deep Neural Network Quantization framework for FPGA
- Fully Quantized Image Super-Resolution Networks
- Analytical aspects of non-differentiable neural networks
- Consensus Based Multi-Layer Perceptrons for Edge Computing
- Min-Max-Plus Neural Networks
- Self-Reorganizing and Rejuvenating CNNs for Increasing Model Capacity Utilization
- GradFreeBits: Gradient Free Bit Allocation for Dynamic Low Precision Neural Networks
- RRNet: Repetition-Reduction Network for Energy Efficient Decoder of Depth Estimation
- Zero-shot Adversarial Quantization
- "BNN - BN = ?": Training Binary Neural Networks without Batch Normalization
- All-You-Can-Fit 8-Bit Flexible Floating-Point Format for Accurate and Memory-Efficient Inference of Deep Neural Networks
- Efficient and Robust Machine Learning for Real-World Systems
- ShortcutFusion: From Tensorflow to FPGA-based accelerator with reuse-aware memory allocation for shortcut data
- DaSGD: Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging
- One Weight Bitwidth to Rule Them All
- Disentangling Neural Architectures and Weights: A Case Study in Supervised Classification
- SoFAr: Shortcut-based Fractal Architectures for Binary Convolutional Neural Networks
- ASCAI: Adaptive Sampling for acquiring Compact AI
- Scaling Neural Network Performance through Customized Hardware Architectures on Reconfigurable Logic
- Quantization Mimic: Towards Very Tiny CNN for Object Detection
- SQWA: Stochastic Quantized Weight Averaging for Improving the Generalization Capability of Low-Precision Deep Neural Networks
- Widening and Squeezing: Towards Accurate and Efficient QNNs
- Training a Binary Weight Object Detector by Knowledge Transfer for Autonomous Driving
- Adaptive Federated Learning With Gradient Compression in Uplink NOMA
- PBGen: Partial Binarization of Deconvolution-Based Generators for Edge Intelligence
- Perturbative GAN: GAN with Perturbation Layers
- APNN-TC: Accelerating Arbitrary Precision Neural Networks on Ampere GPU Tensor Cores
- InstantNet: Automated Generation and Deployment of Instantaneously Switchable-Precision Networks
- Training Multi-bit Quantized and Binarized Networks with A Learnable Symmetric Quantizer
- Iterative Training: Finding Binary Weight Deep Neural Networks with Layer Binarization
- Bayesian Optimized 1-Bit CNNs
- Direct Quantization for Training Highly Accurate Low Bit-width Deep Neural Networks
- Adaptive Precision Training for Resource Constrained Devices
- Adaptive Binary-Ternary Quantization
- DecisiveNets: Training Deep Associative Memories to Solve Complex Machine Learning Problems
- Low-Complexity LSTM Training and Inference with FloatSD8 Weight Representation
- Maximin Optimization for Binary Regression
- RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple Precisions
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Learning Quantized Neural Nets by Coarse Gradient Method for Non-linear Classification
- Sparsity-Control Ternary Weight Networks
- Sum-Rate-Distortion Function for Indirect Multiterminal Source Coding in Federated Learning
- Demystifying and Generalizing BinaryConnect
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural Networks
- Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization
- SmartDeal: Re-Modeling Deep Network Weights for Efficient Inference and Training
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization
- Auto-Split: A General Framework of Collaborative Edge-Cloud AI
- Differentiable Architecture Pruning for Transfer Learning
- Quantized Neural Networks via {-1, +1} Encoding Decomposition and Acceleration
- Weakly Supervised Recovery of Semantic Attributes
- Development of Quantized DNN Library for Exact Hardware Emulation
- DTNN: Energy-efficient Inference with Dendrite Tree Inspired Neural Networks for Edge Vision Applications
- PokeBNN: A Binary Pursuit of Lightweight Accuracy
- Binarized Weight Error Networks With a Transition Regularization Term
- AutoQNN: An End-to-End Framework for Automatically Quantizing Neural Networks
- BitSplit-Net: Multi-bit Deep Neural Network with Bitwise Activation Function
- Phoeni6: a Systematic Approach for Evaluating the Energy Consumption of Neural Networks
- Toward Compact Parameter Representations for Architecture-Agnostic Neural Network Compression
- Robustness-aware 2-bit quantization with real-time performance for neural network
- Self-Adaptive Network Pruning
- BAMSProd: A Step towards Generalizing the Adaptive Optimization Methods to Deep Binary Model
- QuantNet: Learning to Quantize by Learning within Fully Differentiable Framework
- Resource-Efficient Speech Mask Estimation for Multi-Channel Speech Enhancement
- Penetrating the Fog: the Path to Efficient CNN Models
- Full-stack Optimization for Accelerating CNNs with FPGA Validation
- Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy
- Neural Network Activation Quantization with Bitwise Information Bottlenecks
- Accuracy to Throughput Trade-offs for Reduced Precision Neural Networks on Reconfigurable Logic
- Conditional Neural Architecture Search
- Cross-filter compression for CNN inference acceleration