Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
arXiv:1308.3432
Abstract
Stochastic neurons and hard non-linearities can be useful for a number of reasons in deep learning models, but in many cases they pose a challenging problem: how to estimate the gradient of a loss function with respect to the input of such stochastic or non-smooth neurons? I.e., can we "back-propagate" through these stochastic neurons? We examine this question, existing approaches, and compare four families of solutions, applicable in different settings. One of them is the minimum variance unbiased gradient estimator for stochatic binary neurons (a special case of the REINFORCE algorithm). A second approach, introduced here, decomposes the operation of a binary stochastic neuron into a stochastic binary part and a smooth differentiable part, which approximates the expected effect of the pure stochatic binary neuron to first order. A third approach involves the injection of additive or multiplicative noise in a computational graph that is otherwise differentiable. A fourth approach heuristically copies the gradient with respect to the stochastic output directly as an estimator of the gradient with respect to the sigmoid argument (we call this the straight-through estimator). To explore a context where these estimators are useful, we consider a small-scale version of {\em conditional computation}, where sparse stochastic units form a distributed representation of gaters that can turn off in combinatorially many ways large chunks of the computation performed in the rest of the neural network. In this case, it is important that the gating units produce an actual 0 most of the time. The resulting sparsity can be potentially be exploited to greatly reduce the computational cost of large deep networks for which conditional computation would be useful.
arXiv admin note: substantial text overlap with arXiv:1305.2982
References in corpus (4)
Cited by in corpus (653)
- Language Models are Few-Shot Learners
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Normalizing Flows: An Introduction and Review of Current Methods
- Zero-Shot Text-to-Image Generation
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- PACT: Parameterized Clipping Activation for Quantized Neural Networks
- Adversarially Learned Inference
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Binary Neural Networks: A Survey
- Ensemble Distillation for Robust Model Fusion in Federated Learning
- Q8BERT: Quantized 8Bit BERT
- Enabling Spike-based Backpropagation for Training Deep Neural Network Architectures
- Multi-Task Learning with Deep Neural Networks: A Survey
- CogView: Mastering Text-to-Image Generation via Transformers
- Self-Supervised Speech Representation Learning: A Review
- Recent Advances in Autoencoder-Based Representation Learning
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Normalizing Flows for Probabilistic Modeling and Inference
- Structured Pruning for Deep Convolutional Neural Networks: A survey
- Professor Forcing: A New Algorithm for Training Recurrent Networks
- The Heidelberg spiking datasets for the systematic evaluation of spiking neural networks
- Learned Step Size Quantization
- Learning Object Bounding Boxes for 3D Instance Segmentation on Point Clouds
- Model compression via distillation and quantization
- Hierarchical Multiscale Recurrent Neural Networks
- Video Compression With Rate-Distortion Autoencoders
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Disentangled Representation Learning in Cardiac Image Analysis
- A Survey on Methods and Theories of Quantized Neural Networks
- Mixed Precision Quantization of ConvNets via Differentiable Neural Architecture Search
- WRPN: Wide Reduced-Precision Networks
- Learning Discrete Structures for Graph Neural Networks
- Evaluating Stochastic Rankings with Expected Exposure
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference
- Adaptive Extreme Edge Computing for Wearable Devices
- Gradient Estimation Using Stochastic Computation Graphs
- Bringing AI To Edge: From Deep Learning's Perspective
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- The NLP Cookbook: Modern Recipes for Transformer based Deep Learning Architectures
- EfficientFi: Towards Large-Scale Lightweight WiFi Sensing via CSI Compression
- Learning Multimodal Graph-to-Graph Translation for Molecular Optimization
- OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
- Surrogate Gradient Learning in Spiking Neural Networks
- Learning ReLU Networks on Linearly Separable Data: Algorithm, Optimality, and Generalization
- Learning Sparse Neural Networks through Regularization
- Conditional Computation in Neural Networks for faster models
- advertorch v0.1: An Adversarial Robustness Toolbox based on PyTorch
- AdaShare: Learning What To Share For Efficient Deep Multi-Task Learning
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Dream to Control: Learning Behaviors by Latent Imagination
- Learning Factored Representations in a Deep Mixture of Experts
- Neural Network Approximation: Three Hidden Layers Are Enough
- AutoML-Zero: Evolving Machine Learning Algorithms From Scratch
- Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural Networks
- Path-Restore: Learning Network Path Selection for Image Restoration
- Training with Quantization Noise for Extreme Model Compression
- Boundary-Seeking Generative Adversarial Networks
- Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
- Deep AutoRegressive Networks
- Adversarial Generation of Natural Language
- Bottom-up and top-down approaches for the design of neuromorphic processing systems: Tradeoffs and synergies between natural and artificial intelligence
- Whetstone: A Method for Training Deep Artificial Neural Networks for Binary Communication
- Dynamic Model Pruning with Feedback
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- An overview of mixing augmentation methods and augmentation strategies
- Noisy Activation Functions
- Jointly Learning Sentence Embeddings and Syntax with Unsupervised Tree-LSTMs
- Bayesian Bits: Unifying Quantization and Pruning
- Strategic Attentive Writer for Learning Macro-Actions
- Supervised Compression for Resource-Constrained Edge Computing Systems
- On Accurate Evaluation of GANs for Language Generation
- Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch
- MeliusNet: Can Binary Neural Networks Achieve MobileNet-level Accuracy?
- Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models
- Stochastic Thermodynamics of Non-Linear Electronic Circuits: A Realistic Framework for Computing around kT
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision Transformers
- Relation Embedding with Dihedral Group in Knowledge Graph
- Neural Sampling Machine with Stochastic Synapse allows Brain-like Learning and Inference
- Rotated Binary Neural Network
- Latent Weights Do Not Exist: Rethinking Binarized Neural Network Optimization
- AddNet: Deep Neural Networks Using FPGA-Optimized Multipliers
- Neural Speed Reading via Skim-RNN
- Effective Quantization Methods for Recurrent Neural Networks
- Deep Network with Approximation Error Being Reciprocal of Width to Power of Square Root of Depth
- Discovering Neural Wirings
- Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets
- Learning and Querying Fast Generative Models for Reinforcement Learning
- MaskConnect: Connectivity Learning by Gradient Descent
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks
- Bi-GCN: Binary Graph Convolutional Network
- Noise Analysis of Photonic Modulator Neurons
- Trained Quantization Thresholds for Accurate and Efficient Fixed-Point Inference of Deep Neural Networks
- Improving Inference for Neural Image Compression
- HS-GCN: Hamming Spatial Graph Convolutional Networks for Recommendation
- Emergence of Language with Multi-agent Games: Learning to Communicate with Sequences of Symbols
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks
- Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
- Tracking and Mapping in Medical Computer Vision: A Review
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
- Deep Encoder-Decoder Models for Unsupervised Learning of Controllable Speech Synthesis
- Up or Down? Adaptive Rounding for Post-Training Quantization
- Low-Rank Approximations for Conditional Feedforward Computation in Deep Neural Networks
- Natural Language Understanding with Distributed Representation
- Adversarial Examples in Modern Machine Learning: A Review
- PUERT: Probabilistic Under-sampling and Explicable Reconstruction Network for CS-MRI
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- Mixed Precision DNNs: All you need is a good parametrization
- Winning the Lottery with Continuous Sparsification
- QKD: Quantization-aware Knowledge Distillation
- The Free Energy Principle for Perception and Action: A Deep Learning Perspective
- Explaining Question Answering Models through Text Generation
- Universally Quantized Neural Compression
- Back to Simplicity: How to Train Accurate BNNs from Scratch?
- Discovering Low-Precision Networks Close to Full-Precision Networks for Efficient Embedded Inference
- SpotTune: Transfer Learning through Adaptive Fine-tuning
- Parallel Attention Network with Sequence Matching for Video Grounding
- Discrete Variational Autoencoders
- Discrete Flows: Invertible Generative Models of Discrete Data
- Dynamic Neural Networks: A Survey
- DADA: Differentiable Automatic Data Augmentation
- Routing Networks and the Challenges of Modular and Compositional Computation
- Techniques for Learning Binary Stochastic Feedforward Neural Networks
- DVAE++: Discrete Variational Autoencoders with Overlapping Transformations
- Learning Permutations with Sinkhorn Policy Gradient
- Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
- Stochastic Optimization of Sorting Networks via Continuous Relaxations
- MP-DPD: Low-Complexity Mixed-Precision Neural Networks for Energy-Efficient Digital Predistortion of Wideband Power Amplifiers
- Bridging the Accuracy Gap for 2-bit Quantized Neural Networks (QNN)
- IDF++: Analyzing and Improving Integer Discrete Flows for Lossless Compression
- SEALion: a Framework for Neural Network Inference on Encrypted Data
- Knowledge-refined Denoising Network for Robust Recommendation
- NICE: Noise Injection and Clamping Estimation for Neural Network Quantization
- Defensive Quantization: When Efficiency Meets Robustness
- AxTrain: Hardware-Oriented Neural Network Training for Approximate Inference
- Training Binary Neural Networks through Learning with Noisy Supervision
- Convolutional Generative Adversarial Networks with Binary Neurons for Polyphonic Music Generation
- Movement Pruning: Adaptive Sparsity by Fine-Tuning
- Towards Efficient Training for Neural Network Quantization
- PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers
- Accurate and Compact Convolutional Neural Networks with Trained Binarization
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary Activations
- DSConv: Efficient Convolution Operator
- Learning Cross-Scale Weighted Prediction for Efficient Neural Video Compression
- Towards Lossless ANN-SNN Conversion under Ultra-Low Latency with Dual-Phase Optimization
- Overfitting for Fun and Profit: Instance-Adaptive Data Compression
- MuProp: Unbiased Backpropagation for Stochastic Neural Networks
- Zero Time Waste: Recycling Predictions in Early Exit Neural Networks
- CCVS: Context-aware Controllable Video Synthesis
- T-BFA: Targeted Bit-Flip Adversarial Weight Attack
- Scaling Vision with Sparse Mixture of Experts
- Efficient Exact Verification of Binarized Neural Networks
- HashVFL: Defending Against Data Reconstruction Attacks in Vertical Federated Learning
- Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models
- Binary Neural Networks for Memory-Efficient and Effective Visual Place Recognition in Changing Environments
- Noisy Machines: Understanding Noisy Neural Networks and Enhancing Robustness to Analog Hardware Errors Using Distillation
- Learning Sparse Low-Precision Neural Networks With Learnable Regularization
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
- Forward and Backward Information Retention for Accurate Binary Neural Networks
- Exponentially Increasing the Capacity-to-Computation Ratio for Conditional Computation in Deep Learning
- Latency-Aware Differentiable Neural Architecture Search
- MARS: Multi-macro Architecture SRAM CIM-Based Accelerator with Co-designed Compressed Neural Networks
- Neural Network-Optimized Channel Estimator and Training Signal Design for MIMO Systems with Few-Bit ADCs
- Predicting Video with VQVAE
- Attention over Parameters for Dialogue Systems
- Learning to Segment Inputs for NMT Favors Character-Level Processing
- AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search
- DeepSym: Deep Symbol Generation and Rule Learning from Unsupervised Continuous Robot Interaction for Planning
- Hierarchical Autoregressive Image Models with Auxiliary Decoders
- ProbAct: A Probabilistic Activation Function for Deep Neural Networks
- Mastering Atari with Discrete World Models
- End-to-End Supermask Pruning: Learning to Prune Image Captioning Models
- Neural Language Generation: Formulation, Methods, and Evaluation
- Efficient Certified Defenses Against Patch Attacks on Image Classifiers
- Self-Binarizing Networks
- Polygonal Building Segmentation by Frame Field Learning
- Conducting Credit Assignment by Aligning Local Representations
- Lossy Image Compression with Normalizing Flows
- Differentiable JPEG: The Devil is in the Details
- Attacking Binarized Neural Networks
- Towards Effective Low-bitwidth Convolutional Neural Networks
- SEQ^3: Differentiable Sequence-to-Sequence-to-Sequence Autoencoder for Unsupervised Abstractive Sentence Compression
- SNN2ANN: A Fast and Memory-Efficient Training Framework for Spiking Neural Networks
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action Recognition
- Fighting Quantization Bias With Bias
- Regularizing Activation Distribution for Training Binarized Deep Networks
- Degree-Quant: Quantization-Aware Training for Graph Neural Networks
- Resource-Efficient Neural Networks for Embedded Systems
- Differentiable Model Compression via Pseudo Quantization Noise
- Plug & Play Directed Evolution of Proteins with Gradient-based Discrete MCMC
- Taming Visually Guided Sound Generation
- Modeling Lost Information in Lossy Image Compression
- A RAD approach to deep mixture models
- Role-Wise Data Augmentation for Knowledge Distillation
- BlockCopy: High-Resolution Video Processing with Block-Sparse Feature Propagation and Online Policies
- Deep Directed Generative Autoencoders
- SegBlocks: Block-Based Dynamic Resolution Networks for Real-Time Segmentation
- -ARM: Network Sparsification via Stochastic Binary Optimization
- Self-Supervised Exploration via Disagreement
- Deep Learning as a Mixed Convex-Combinatorial Optimization Problem
- Self-Distribution Binary Neural Networks
- Few Shot Network Compression via Cross Distillation
- LSQ+: Improving low-bit quantization through learnable offsets and better initialization
- Neural Feature Search for RGB-Infrared Person Re-Identification
- Robust Lottery Tickets for Pre-trained Language Models
- Pruning Self-attentions into Convolutional Layers in Single Path
- Training Binary Neural Networks using the Bayesian Learning Rule
- Bit-Flip Attack: Crushing Neural Network with Progressive Bit Search
- Generating Diverse and Meaningful Captions
- Estimating Gradients for Discrete Random Variables by Sampling without Replacement
- Switchable Precision Neural Networks
- A Tandem Learning Rule for Effective Training and Rapid Inference of Deep Spiking Neural Networks
- A Review of Learning with Deep Generative Models from Perspective of Graphical Modeling
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
- Automatic Mixed-Precision Quantization Search of BERT
- Relaxed Quantization for Discretized Neural Networks
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- Towards the AlexNet Moment for Homomorphic Encryption: HCNN, theFirst Homomorphic CNN on Encrypted Data with GPUs
- Biased Mixtures Of Experts: Enabling Computer Vision Inference Under Data Transfer Limitations
- Efficient Neural Network Training via Forward and Backward Propagation Sparsification
- On Quantizing Implicit Neural Representations
- Adversarial Defense via Data Dependent Activation Function and Total Variation Minimization
- Binary Ensemble Neural Network: More Bits per Network or More Networks per Bit?
- Dynamic Deep Neural Networks: Optimizing Accuracy-Efficiency Trade-offs by Selective Execution
- Additive Noise Annealing and Approximation Properties of Quantized Neural Networks
- Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N:M Transposable Masks
- BitPruning: Learning Bitlengths for Aggressive and Accurate Quantization
- Learning to Drop: Robust Graph Neural Network via Topological Denoising
- Learning Neural-Symbolic Descriptive Planning Models via Cube-Space Priors: The Voyage Home (to STRIPS)
- BatchQuant: Quantized-for-all Architecture Search with Robust Quantizer
- PixelVAE++: Improved PixelVAE with Discrete Prior
- Tangent: Automatic Differentiation Using Source Code Transformation in Python
- Early Exiting with Ensemble Internal Classifiers
- Latent Tree Learning with Differentiable Parsers: Shift-Reduce Parsing and Chart Parsing
- A Statistical Framework for Low-bitwidth Training of Deep Neural Networks
- Single-pass Object-adaptive Data Undersampling and Reconstruction for MRI
- EDUCE: Explaining model Decisions through Unsupervised Concepts Extraction
- Stochastic Generative Hashing
- Hierarchical Training of Deep Neural Networks Using Early Exiting
- KDLSQ-BERT: A Quantized Bert Combining Knowledge Distillation with Learned Step Size Quantization
- A White Paper on Neural Network Quantization
- Eliminating Exposure Bias and Loss-Evaluation Mismatch in Multiple Object Tracking
- One Loss for All: Deep Hashing with a Single Cosine Similarity based Learning Objective
- Learning Navigation Subroutines from Egocentric Videos
- Efficient Bitwidth Search for Practical Mixed Precision Neural Network
- NeuroPack: An Algorithm-level Python-based Simulator for Memristor-empowered Neuro-inspired Computing
- Learn to Match: Automatic Matching Network Design for Visual Tracking
- FBNetV5: Neural Architecture Search for Multiple Tasks in One Run
- Controlling Computation versus Quality for Neural Sequence Models
- Towards More Human-like AI Communication: A Review of Emergent Communication Research
- Learning to Hash with Graph Neural Networks for Recommender Systems
- Blended Coarse Gradient Descent for Full Quantization of Deep Neural Networks
- Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight Masks
- Multi-Prize Lottery Ticket Hypothesis: Finding Accurate Binary Neural Networks by Pruning A Randomly Weighted Network
- Contrastive Learning for Debiased Candidate Generation in Large-Scale Recommender Systems
- Pairwise Supervised Hashing with Bernoulli Variational Auto-Encoder and Self-Control Gradient Estimator
- Contextual Dropout: An Efficient Sample-Dependent Dropout Module
- Vector representations of text data in deep learning
- Training Sparse Neural Networks
- Revisiting the Hierarchical Multiscale LSTM
- Learning K-way D-dimensional Discrete Code For Compact Embedding Representations
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
- Compression of Acoustic Event Detection Models with Low-rank Matrix Factorization and Quantization Training
- ECM: Early Exit via Class Means for Efficient Supervised and Unsupervised Learning
- A Lite Distributed Semantic Communication System for Internet of Things
- Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE
- Gradient Regularization for Quantization Robustness
- Improving Gradient Estimation in Evolutionary Strategies With Past Descent Directions
- CAT: Compression-Aware Training for bandwidth reduction
- AR-Net: Adaptive Frame Resolution for Efficient Action Recognition
- On-FPGA Training with Ultra Memory Reduction: A Low-Precision Tensor Method
- Variable-rate discrete representation learning
- A Brief Introduction to Generative Models
- Modelling Latent Translations for Cross-Lingual Transfer
- Attention-Based Generative Neural Image Compression on Solar Dynamics Observatory
- Instance-Adaptive Video Compression: Improving Neural Codecs by Training on the Test Set
- BiPointNet: Binary Neural Network for Point Clouds
- Efficient Marginalization of Discrete and Structured Latent Variables via Sparsity
- Training High-Performance and Large-Scale Deep Neural Networks with Full 8-bit Integers
- Efficient Deep Neural Networks
- Integer Discrete Flows and Lossless Compression
- SQuantizer: Simultaneous Learning for Both Sparse and Low-precision Neural Networks
- Any-Precision Deep Neural Networks
- Layer-wise Learning of Stochastic Neural Networks with Information Bottleneck
- Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
- Learning Task-Specific Strategies for Accelerated MRI
- Recognition-Aware Learned Image Compression
- SenSeNet: Neural Keyphrase Generation with Document Structure
- Dynamic Frame Interpolation in Wavelet Domain
- Training Compact Neural Networks with Binary Weights and Low Precision Activations
- Modularity in Deep Learning: A Survey
- Espresso: Efficient Forward Propagation for BCNNs
- Learning Frequency Domain Approximation for Binary Neural Networks
- Per-Tensor Fixed-Point Quantization of the Back-Propagation Algorithm
- Learning Visual Question Answering by Bootstrapping Hard Attention
- Dynamics-aware Adversarial Attack of Adaptive Neural Networks
- DisARM: An Antithetic Gradient Estimator for Binary Latent Variables
- Triple Generative Adversarial Networks
- OMPQ: Orthogonal Mixed Precision Quantization
- Boosting Binary Masks for Multi-Domain Learning through Affine Transformations
- Dynamic Multi-Branch Layers for On-Device Neural Machine Translation
- InfoCNF: An Efficient Conditional Continuous Normalizing Flow with Adaptive Solvers
- Reparameterization trick for discrete variables
- Learning Discrete Structured Representations by Adversarially Maximizing Mutual Information
- VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
- Scalable Model Compression by Entropy Penalized Reparameterization
- Extending LOUPE for K-space Under-sampling Pattern Optimization in Multi-coil MRI
- Improving Adversarial Robustness in Weight-quantized Neural Networks
- A Latent Morphology Model for Open-Vocabulary Neural Machine Translation
- An Integer Programming Approach to Deep Neural Networks with Binary Activation Functions
- Dynamic Capacity Networks
- ENGINE: Energy-Based Inference Networks for Non-Autoregressive Machine Translation
- Predicting distributions with Linearizing Belief Networks
- Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
- AnalogNets: ML-HW Co-Design of Noise-robust TinyML Models and Always-On Analog Compute-in-Memory Accelerator
- Sharpness-aware Quantization for Deep Neural Networks
- FLightNNs: Lightweight Quantized Deep Neural Networks for Fast and Accurate Inference
- Simultaneously Optimizing Weight and Quantizer of Ternary Neural Network using Truncated Gaussian Approximation
- ACtuAL: Actor-Critic Under Adversarial Learning
- Network Quantization with Element-wise Gradient Scaling
- Learning Versatile Convolution Filters for Efficient Visual Recognition
- Training Generative Adversarial Networks with Binary Neurons by End-to-end Backpropagation
- Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
- Mix and Match: A Novel FPGA-Centric Deep Neural Network Quantization Framework
- Deep Spiking Neural Networks for Large Vocabulary Automatic Speech Recognition
- The Variational Bandwidth Bottleneck: Stochastic Evaluation on an Information Budget
- The Bayesian Learning Rule
- Effective and Fast: A Novel Sequential Single Path Search for Mixed-Precision Quantization
- Propagating Asymptotic-Estimated Gradients for Low Bitwidth Quantized Neural Networks
- MWQ: Multiscale Wavelet Quantized Neural Networks
- Soft then Hard: Rethinking the Quantization in Neural Image Compression
- Inducing Interpretable Representations with Variational Autoencoders
- Neural Hybrid Automata: Learning Dynamics with Multiple Modes and Stochastic Transitions
- Improving Deep Representation Learning via Auxiliary Learnable Target Coding
- Towards Lossless Binary Convolutional Neural Networks Using Piecewise Approximation
- Cascaded channel pruning using hierarchical self-distillation
- Unsupervised Hashing with Contrastive Information Bottleneck
- Efficient Halftoning via Deep Reinforcement Learning
- Weight Pruning via Adaptive Sparsity Loss
- Conditional Computation for Continual Learning
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation
- Binarizing by Classification: Is soft function really necessary?
- Transferable Sparse Adversarial Attack
- Hierarchical Autoregressive Modeling for Neural Video Compression
- Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data
- Taming Binarized Neural Networks and Mixed-Integer Programs
- SpVOS: Efficient Video Object Segmentation with Triple Sparse Convolution
- Anytime Sampling for Autoregressive Models via Ordered Autoencoding
- Meta Approach to Data Augmentation Optimization
- Generative Models for Improved Naturalness, Intelligibility, and Voicing of Whispered Speech
- Differentiable Dynamic Quantization with Mixed Precision and Adaptive Resolution
- Accelerating SNN Training with Stochastic Parallelizable Spiking Neurons
- Towards Interpretable and Reliable Reading Comprehension: A Pipeline Model with Unanswerability Prediction
- Hide-and-Seek: A Template for Explainable AI
- Divide and Conquer: Leveraging Intermediate Feature Representations for Quantized Training of Neural Networks
- End-to-End Supervised Product Quantization for Image Search and Retrieval
- Implicit MLE: Backpropagating Through Discrete Exponential Family Distributions
- Algorithm and Hardware Design of Discrete-Time Spiking Neural Networks Based on Back Propagation with Binary Activations
- You Look Twice: GaterNet for Dynamic Filter Selection in CNNs
- Improved Gradient-Based Optimization Over Discrete Distributions
- A Brief Overview of Unsupervised Neural Speech Representation Learning
- End-to-End Feedback Loss in Speech Chain Framework via Straight-Through Estimator
- Leap-LSTM: Enhancing Long Short-Term Memory for Text Categorization
- A 1Mb mixed-precision quantized encoder for image classification and patch-based compression
- Measuring the Biases and Effectiveness of Content-Style Disentanglement
- BitNet: Bit-Regularized Deep Neural Networks
- Low-Power Computer Vision: Status, Challenges, Opportunities
- Hierarchical Multi-scale Attention Networks for Action Recognition
- Surprisal-Triggered Conditional Computation with Neural Networks
- VA-RED: Video Adaptive Redundancy Reduction
- Mixed-Precision Quantized Neural Network with Progressively Decreasing Bitwidth For Image Classification and Object Detection
- Improving Accuracy of Binary Neural Networks using Unbalanced Activation Distribution
- Accelerating Learnt Video Codecs with Gradient Decay and Layer-wise Distillation
- Differentiable Joint Pruning and Quantization for Hardware Efficiency
- Straight-Through Estimator as Projected Wasserstein Gradient Flow
- Mastering emergent language: learning to guide in simulated navigation
- Exploring the Back Alleys: Analysing The Robustness of Alternative Neural Network Architectures against Adversarial Attacks
- Embarrassingly Simple Binary Representation Learning
- Towards Unsupervised Language Understanding and Generation by Joint Dual Learning
- Auto-Encoding Twin-Bottleneck Hashing
- mcLARO: Multi-Contrast Learned Acquisition and Reconstruction Optimization for simultaneous quantitative multi-parametric mapping
- Adaptive Precision Training (AdaPT): A dynamic fixed point quantized training approach for DNNs
- One Self-Configurable Model to Solve Many Abstract Visual Reasoning Problems
- No Multiplication? No Floating Point? No Problem! Training Networks for Efficient Inference
- Reinforcement Learning with Feedback-modulated TD-STDP
- Generative Semantic Hashing Enhanced via Boltzmann Machines
- Single-Path Mobile AutoML: Efficient ConvNet Design and NAS Hyperparameter Optimization
- Inference with Hybrid Bio-hardware Neural Networks
- Dichotomize and Generalize: PAC-Bayesian Binary Activated Deep Neural Networks
- A High-Performance Adaptive Quantization Approach for Edge CNN Applications
- Language as a Latent Sequence: deep latent variable models for semi-supervised paraphrase generation
- Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text Recognition
- Simplified Stochastic Feedforward Neural Networks
- Pelta: Shielding Transformers to Mitigate Evasion Attacks in Federated Learning
- Predicting Adversarial Examples with High Confidence
- Energy Confused Adversarial Metric Learning for Zero-Shot Image Retrieval and Clustering
- Effective Sparsification of Neural Networks with Global Sparsity Constraint
- Learnable Companding Quantization for Accurate Low-bit Neural Networks
- Index Tracking with Cardinality Constraints: A Stochastic Neural Networks Approach
- Direct Evolutionary Optimization of Variational Autoencoders With Binary Latents
- Dynamically Throttleable Neural Networks (TNN)
- MSP: An FPGA-Specific Mixed-Scheme, Multi-Precision Deep Neural Network Quantization Framework
- A Probabilistic End-To-End Task-Oriented Dialog Model with Latent Belief States towards Semi-Supervised Learning
- Perceive, Attend, and Drive: Learning Spatial Attention for Safe Self-Driving
- Low-Rank Training of Deep Neural Networks for Emerging Memory Technology
- Alternating Direction Method of Multipliers for Quantization
- Generalized Ternary Connect: End-to-End Learning and Compression of Multiplication-Free Deep Neural Networks
- Efficient non-uniform quantizer for quantized neural network targeting reconfigurable hardware
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network Quantization
- FAT: Learning Low-Bitwidth Parametric Representation via Frequency-Aware Transformation
- Critical initialisation in continuous approximations of binary neural networks
- Integrating Semantics and Neighborhood Information with Graph-Driven Generative Models for Document Retrieval
- Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies
- Backprop-Q: Generalized Backpropagation for Stochastic Computation Graphs
- Neural Machine Translation with 4-Bit Precision and Beyond
- Improving Binary Neural Networks through Fully Utilizing Latent Weights
- Training Deep Spiking Auto-encoders without Bursting or Dying Neurons through Regularization
- Efficient Neural Architecture Search for End-to-end Speech Recognition via Straight-Through Gradients
- Adversarial Contrastive Pre-training for Protein Sequences
- Information contraction in noisy binary neural networks and its implications
- One Timestep is All You Need: Training Spiking Neural Networks with Ultra Low Latency
- Continual Learning via Bit-Level Information Preserving
- Confounding Tradeoffs for Neural Network Quantization
- Hierarchical Autoencoder-based Lossy Compression for Large-scale High-resolution Scientific Data
- Distribution-sensitive Information Retention for Accurate Binary Neural Network
- NASB: Neural Architecture Search for Binary Convolutional Neural Networks
- Automatic Validation of Textual Attribute Values in E-commerce Catalog by Learning with Limited Labeled Data
- Emergence of Numeric Concepts in Multi-Agent Autonomous Communication
- Reducing the Computational Cost of Deep Generative Models with Binary Neural Networks
- Image Captioning with Sparse Recurrent Neural Network
- Dynamic Collective Intelligence Learning: Finding Efficient Sparse Model via Refined Gradients for Pruned Weights
- Intelligence plays dice: Stochasticity is essential for machine learning
- Quantized Convolutional Neural Networks Through the Lens of Partial Differential Equations
- Weakly Supervised Concept Map Generation through Task-Guided Graph Translation
- Iterative Refinement of the Approximate Posterior for Directed Belief Networks
- High-Capacity Expert Binary Networks
- Neural-Symbolic Descriptive Action Model from Images: The Search for STRIPS
- Drawing Robust Scratch Tickets: Subnetworks with Inborn Robustness Are Found within Randomly Initialized Networks
- Sampling-Free Learning of Bayesian Quantized Neural Networks
- Progressive Learning of Low-Precision Networks
- Dynamic Network Quantization for Efficient Video Inference
- Neurocoder: Learning General-Purpose Computation Using Stored Neural Programs
- Image-to-Image Translation with Low Resolution Conditioning
- Histogram-Equalized Quantization for logic-gated Residual Neural Networks
- Learning Sparse Mixture of Experts for Visual Question Answering
- FrostNet: Towards Quantization-Aware Network Architecture Search
- Probabilistic Mixture-of-Experts for Efficient Deep Reinforcement Learning
- Path Sample-Analytic Gradient Estimators for Stochastic Binary Networks
- On the Effects of Quantisation on Model Uncertainty in Bayesian Neural Networks
- Invertible Image Rescaling
- Unsupervised Acoustic Unit Representation Learning for Voice Conversion using WaveNet Auto-encoders
- Semi-supervised Learning for Multi-speaker Text-to-speech Synthesis Using Discrete Speech Representation
- Question Guided Modular Routing Networks for Visual Question Answering
- crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder
- Document Hashing with Mixture-Prior Generative Models
- WSRNet: Joint Spotting and Recognition of Handwritten Words
- Matching-oriented Product Quantization For Ad-hoc Retrieval
- Synthetic Aperture Radar Image Change Detection via Siamese Adaptive Fusion Network
- Learned Variable-Rate Image Compression with Residual Divisive Normalization
- ReCU: Reviving the Dead Weights in Binary Neural Networks
- Binarized Canonical Polyadic Decomposition for Knowledge Graph Completion
- Learning Accurate Decision Trees with Bandit Feedback via Quantized Gradient Descent
- SiMaN: Sign-to-Magnitude Network Binarization
- DNN Feature Map Compression using Learned Representation over GF(2)
- Efficient Mixed Precision Quantization in Graph Neural Networks
- RTN: Reparameterized Ternary Network
- Unlocking Pixels for Reinforcement Learning via Implicit Attention
- Slot Machines: Discovering Winning Combinations of Random Weights in Neural Networks
- Closing the Dequantization Gap: PixelCNN as a Single-Layer Flow
- Population-based Gradient Descent Weight Learning for Graph Coloring Problems
- Training Quantized Neural Networks with a Full-precision Auxiliary Module
- COMET: Learning Cardinality Constrained Mixture of Experts with Trees and Local Search
- Searching for Accurate Binary Neural Architectures
- BWCP: Probabilistic Learning-to-Prune Channels for ConvNets via Batch Whitening
- Fast Adjustable Threshold For Uniform Neural Network Quantization (Winning solution of LPIRC-II)
- Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient Estimator
- BILLNET: A Binarized Conv3D-LSTM Network with Logic-gated residual architecture for hardware-efficient video inference
- Prototypical Contrastive Learning and Adaptive Interest Selection for Candidate Generation in Recommendations
- Improving Efficiency in Neural Network Accelerator Using Operands Hamming Distance optimization
- NASH: Toward End-to-End Neural Architecture for Generative Semantic Hashing
- Mask-aware networks for crowd counting
- Invertible Gaussian Reparameterization: Revisiting the Gumbel-Softmax
- Inherent Weight Normalization in Stochastic Neural Networks
- Hierarchical Photo-Scene Encoder for Album Storytelling
- Binarized Knowledge Graph Embeddings
- NWT: Towards natural audio-to-video generation with representation learning
- KCNet: An Insect-Inspired Single-Hidden-Layer Neural Network with Randomized Binary Weights for Prediction and Classification Tasks
- Optimal Variance Control of the Score Function Gradient Estimator for Importance Weighted Bounds
- AQD: Towards Accurate Fully-Quantized Object Detection
- Variational Latent-State GPT for Semi-Supervised Task-Oriented Dialog Systems
- Deep Learning for Distributed Channel Feedback and Multiuser Precoding in FDD Massive MIMO
- Continual Learning: Forget-free Winning Subnetworks for Video Representations
- Causal Attention for Unbiased Visual Recognition
- Universal Approximation of Functions on Sets
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- On (Emergent) Systematic Generalisation and Compositionality in Visual Referential Games with Straight-Through Gumbel-Softmax Estimator
- Learning Sampling in Financial Statement Audits using Vector Quantised Autoencoder Neural Networks
- Generating Semantically Valid Adversarial Questions for TableQA
- WaveQ: Gradient-Based Deep Quantization of Neural Networks through Sinusoidal Adaptive Regularization
- Foveation for Segmentation of Ultra-High Resolution Images
- Wasserstein Routed Capsule Networks
- Interpretable Neural Network Decoupling
- Binary Stochastic Filtering: feature selection and beyond
- Quantum Annealing Formulation for Binary Neural Networks
- One Weight Bitwidth to Rule Them All
- The Image Local Autoregressive Transformer
- Accelerator-Aware Training for Transducer-Based Speech Recognition
- Combinatorial Optimization for Panoptic Segmentation: A Fully Differentiable Approach
- Hessian-aware Quantized Node Embeddings for Recommendation
- MOGNET: A Mux-residual quantized Network leveraging Online-Generated weights
- Development of systematic uncertainty-aware neural network trainings for binned-likelihood analyses at the LHC
- Sub-quadratic scalable approximate linear converter using multi-plane light conversion with low-entropy mode mixers
- SCoTTi: Save Computation at Training Time with an adaptive framework
- Taming Reversible Halftoning via Predictive Luminance
- Abstract Reasoning via Logic-guided Generation
- Learning Effective and Efficient Embedding via an Adaptively-Masked Twins-based Layer
- HERO: Hessian-Enhanced Robust Optimization for Unifying and Improving Generalization and Quantization Performance
- Math Word Problem Generation with Mathematical Consistency and Problem Context Constraints
- Sparse MoEs meet Efficient Ensembles
- Soft Actor-Critic With Integer Actions
- Rapid Elastic Architecture Search under Specialized Classes and Resource Constraints
- Pruning Ternary Quantization
- Fitting summary statistics of neural data with a differentiable spiking network simulator
- 3U-EdgeAI: Ultra-Low Memory Training, Ultra-Low BitwidthQuantization, and Ultra-Low Latency Acceleration
- Model-Agnostic Meta-Attack: Towards Reliable Evaluation of Adversarial Robustness
- Improved Techniques for Quantizing Deep Networks with Adaptive Bit-Widths
- Part & Whole Extraction: Towards A Deep Understanding of Quantitative Facts for Percentages in Text
- CARMS: Categorical-Antithetic-REINFORCE Multi-Sample Gradient Estimator
- Fully Quantized Image Super-Resolution Networks
- Learning Discrete Energy-based Models via Auxiliary-variable Local Exploration
- Neuro-Symbolic Representations for Video Captioning: A Case for Leveraging Inductive Biases for Vision and Language
- Training of mixed-signal optical convolutional neural network with reduced quantization level
- FATNN: Fast and Accurate Ternary Neural Networks
- Knowledge Distillation-aided End-to-End Learning for Linear Precoding in Multiuser MIMO Downlink Systems with Finite-Rate Feedback
- DiffPrune: Neural Network Pruning with Deterministic Approximate Binary Gates and Regularization
- Composed Fine-Tuning: Freezing Pre-Trained Denoising Autoencoders for Improved Generalization
- BiSNN: Training Spiking Neural Networks with Binary Weights via Bayesian Learning
- CoDeNet: Efficient Deployment of Input-Adaptive Object Detection on Embedded FPGAs
- Reliable Categorical Variational Inference with Mixture of Discrete Normalizing Flows
- Towards Recognizing New Semantic Concepts in New Visual Domains
- End-to-end Generative Zero-shot Learning via Few-shot Learning
- GradFreeBits: Gradient Free Bit Allocation for Dynamic Low Precision Neural Networks
- Scalable Approximate Inference and Some Applications
- Dynamic Narrowing of VAE Bottlenecks Using GECO and L0 Regularization
- Temporal Feature Fusion with Sampling Pattern Optimization for Multi-echo Gradient Echo Acquisition and Image Reconstruction
- Learning Multi-granular Quantized Embeddings for Large-Vocab Categorical Features in Recommender Systems
- "BNN - BN = ?": Training Binary Neural Networks without Batch Normalization
- Playing log(N)-Questions over Sentences
- Learn to Compress CSI and Allocate Resources in Vehicular Networks
- Cooperative image captioning
- SoFAr: Shortcut-based Fractal Architectures for Binary Convolutional Neural Networks
- Binary Stochastic Filtering: a Method for Neural Network Size Minimization and Supervised Feature Selection
- Efficient and Robust Machine Learning for Real-World Systems
- Supervised Vector Quantized Variational Autoencoder for Learning Interpretable Global Representations
- Learning to Request Guidance in Emergent Communication
- Deep Learning-based Image Compression with Trellis Coded Quantization
- Depth-Adaptive Graph Recurrent Network for Text Classification
- A Main/Subsidiary Network Framework for Simplifying Binary Neural Network
- Recurrent Binary Embedding for GPU-Enabled Exhaustive Retrieval from Billion-Scale Semantic Vectors
- Deep Hashing using Entropy Regularised Product Quantisation Network
- Incorporating Symbolic Sequential Modeling for Speech Enhancement
- Memory-Augmented Temporal Dynamic Learning for Action Recognition
- Differentiable Sampling with Flexible Reference Word Order for Neural Machine Translation
- Toward Runtime-Throttleable Neural Networks
- Depthwise Discrete Representation Learning
- QuantNet: Learning to Quantize by Learning within Fully Differentiable Framework
- Binary Stochastic Representations for Large Multi-class Classification
- Deep Model Compression Via Two-Stage Deep Reinforcement Learning
- Compression of Acoustic Event Detection Models With Quantized Distillation
- ENOS: Energy-Aware Network Operator Search for Hybrid Digital and Compute-in-Memory DNN Accelerators
- Notes on Latent Structure Models and SPIGOT
- Channel selection using Gumbel Softmax
- Generalizing Emergent Communication
- Discrete Variational Attention Models for Language Generation
- BAMSProd: A Step towards Generalizing the Adaptive Optimization Methods to Deep Binary Model
- Learned Multi-Resolution Variable-Rate Image Compression with Octave-based Residual Blocks
- Compressing Deep Convolutional Neural Networks by Stacking Low-dimensional Binary Convolution Filters
- Memory and Computation-Efficient Kernel SVM via Binary Embedding and Ternary Model Coefficients
- Explainability as statistical inference
- Adversarial Multi-scale Feature Learning for Person Re-identification
- Learning to Localize Through Compressed Binary Maps
- FIVES: Feature Interaction Via Edge Search for Large-Scale Tabular Data
- Deconstructing the Structure of Sparse Neural Networks
- Bi-Directional Differentiable Input Reconstruction for Low-Resource Neural Machine Translation
- TaylorGAN: Neighbor-Augmented Policy Update for Sample-Efficient Natural Language Generation
- Understanding the wiring evolution in differentiable neural architecture search
- Filter Pre-Pruning for Improved Fine-tuning of Quantized Deep Neural Networks
- A2: Extracting Cyclic Switchings from DOB-nets for Rejecting Excessive Disturbances
- Hindsight Network Credit Assignment
- Learning to Reconstruct and Segment 3D Objects
- Learning Quantized Neural Nets by Coarse Gradient Method for Non-linear Classification
- Net: Augmented Parallel-Pyramid Net for Attention Guided Pose Estimation
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- RMSMP: A Novel Deep Neural Network Quantization Framework with Row-wise Mixed Schemes and Multiple Precisions
- Fast and Flexible Image Blind Denoising via Competition of Experts
- Demystifying and Generalizing BinaryConnect
- A Fourier View of REINFORCE
- Multitask Adaptation by Retrospective Exploration with Learned World Models
- Weakly Supervised Recovery of Semantic Attributes
- CBP: Backpropagation with constraint on weight precision using a pseudo-Lagrange multiplier method
- Efficient and Robust Mixed-Integer Optimization Methods for Training Binarized Deep Neural Networks
- BERMo: What can BERT learn from ELMo?
- Disentangling Representations of Text by Masking Transformers
- Plan, Attend, Generate: Character-level Neural Machine Translation with Planning in the Decoder
- Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural Networks
- Discrete representations in neural models of spoken language
- Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization
- Invertible Tone Mapping with Selectable Styles
- Variational latent discrete representation for time series modelling
- Minimizing Communication while Maximizing Performance in Multi-Agent Reinforcement Learning
- CHISEL: Compression-Aware High-Accuracy Embedded Indoor Localization with Deep Learning
- Exact Backpropagation in Binary Weighted Networks with Group Weight Transformations
- Unsupervised Word Segmentation from Discrete Speech Units in Low-Resource Settings
- On Anytime Learning at Macroscale
- Learning Multi-Layered GBDT Via Back Propagation
- Meta-Aggregator: Learning to Aggregate for 1-bit Graph Neural Networks
- : Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift Networks
- Breaking the Conventional Forward-Backward Tie in Neural Networks: Activation Functions
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization
- Continual learning under domain transfer with sparse synaptic bursting
- Learning to Extend Program Graphs to Work-in-Progress Code
- BitSplit-Net: Multi-bit Deep Neural Network with Bitwise Activation Function
- Quantized Neural Networks via {-1, +1} Encoding Decomposition and Acceleration
- Unbiased Gradient Estimation with Balanced Assignments for Mixtures of Experts
- StreamHover: Livestream Transcript Summarization and Annotation
- Refining BERT Embeddings for Document Hashing via Mutual Information Maximization
- Representation Learning for Efficient and Effective Similarity Search and Recommendation
- Discrete Auto-regressive Variational Attention Models for Text Modeling
- Discrete Tree Flows via Tree-Structured Permutations
- 4-bit Quantization of LSTM-based Speech Recognition Models
- A Survey on Green Deep Learning
- Learning Energy-Based Approximate Inference Networks for Structured Applications in NLP
- BOBCAT: Bilevel Optimization-Based Computerized Adaptive Testing
- LLC: Accurate, Multi-purpose Learnt Low-dimensional Binary Codes
- Memory-Efficient Factorization Machines via Binarizing both Data and Model Coefficients
- Quantization Loss Re-Learning Method
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Solving Large-Scale 0-1 Knapsack Problems and its Application to Point Cloud Resampling
- Cross-modal Spectrum Transformation Network For Acoustic Scene classification
- Model-Driven Deep Learning for Massive MU-MIMO with Finite-Alphabet Precoding
- Towards Modality Transferable Visual Information Representation with Optimal Model Compression
- Lattice Representation Learning
- Learn to Allocate Resources in Vehicular Networks
- Evidence-Aware Inferential Text Generation with Vector Quantised Variational AutoEncoder
- Neural Plasticity Networks
- Obfuscation for Privacy-preserving Syntactic Parsing
- PR Product: A Substitute for Inner Product in Neural Networks
- Latent Transformations for Discrete-Data Normalising Flows
- Minimalist Regression Network with Reinforced Gradients and Weighted Estimates: a Case Study on Parameters Estimation in Automated Welding
- Differentiable TAN Structure Learning for Bayesian Network Classifiers
- Target Propagation via Regularized Inversion
- Semi-Relaxed Quantization with DropBits: Training Low-Bit Neural Networks via Bit-wise Regularization
- Automatic low-bit hybrid quantization of neural networks through meta learning