Do Deep Nets Really Need to be Deep?
arXiv:1312.6184
Abstract
Currently, deep neural networks are the state of the art on problems such as speech recognition and computer vision. In this extended abstract, we show that shallow feed-forward networks can learn the complex functions previously learned by deep nets and achieve accuracies previously only achievable with deep models. Moreover, in some cases the shallow neural nets can learn these deep functions using a total number of parameters similar to the original deep model. We evaluate our method on the TIMIT phoneme recognition task and are able to train shallow fully-connected nets that perform similarly to complex, well-engineered, deep convolutional architectures. Our success in training shallow neural nets to mimic deeper models suggests that there probably exist better algorithms for training shallow feed-forward nets than those currently available.
final revision coming soon
References in corpus (3)
Cited by in corpus (400)
- Knowledge Distillation: A Survey
- FitNets: Hints for Thin Deep Nets
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- FractalNet: Ultra-Deep Neural Networks without Residuals
- Tensorizing Neural Networks
- Impact of Fully Connected Layers on Performance of Convolutional Neural Networks for Image Classification
- Compressing Neural Networks with the Hashing Trick
- Born Again Neural Networks
- Knowledge Distillation by On-the-Fly Native Ensemble
- Hello Edge: Keyword Spotting on Microcontrollers
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Bayesian Compression for Deep Learning
- A Review on Deep Learning Techniques for Video Prediction
- Split Computing and Early Exiting for Deep Learning Applications: Survey and Research Challenges
- Self-training with Noisy Student improves ImageNet classification
- SoundNet: Learning Sound Representations from Unlabeled Video
- Adaptive Multi-Teacher Multi-level Knowledge Distillation
- Adversarial Examples: Attacks and Defenses for Deep Learning
- Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead
- Data-Free Knowledge Distillation for Deep Neural Networks
- DocBERT: BERT for Document Classification
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
- Group Knowledge Transfer: Federated Learning of Large CNNs at the Edge
- Verifiable Reinforcement Learning via Policy Extraction
- Unifying distillation and privileged information
- Anomaly Detection in Univariate Time-series: A Survey on the State-of-the-Art
- Distilled Siamese Networks for Visual Tracking
- A Practical Survey on Faster and Lighter Transformers
- Crypto-Nets: Neural Networks over Encrypted Data
- Distill-and-Compare: Auditing Black-Box Models Using Transparent Model Distillation
- N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning
- SECNLP: A Survey of Embeddings in Clinical Natural Language Processing
- Distilling Knowledge from Deep Networks with Applications to Healthcare Domain
- Compressing Recurrent Neural Network with Tensor Train
- Light-Weight RefineNet for Real-Time Semantic Segmentation
- Recurrent Neural Network Training with Dark Knowledge Transfer
- Deep learning generalizes because the parameter-function map is biased towards simple functions
- Sequence-Level Knowledge Distillation
- Review: Deep Learning in Electron Microscopy
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Dataset Condensation with Gradient Matching
- Do Deep Convolutional Nets Really Need to be Deep and Convolutional?
- A Teacher-Student Framework for Zero-Resource Neural Machine Translation
- Deep Model Compression: Distilling Knowledge from Noisy Teachers
- ISeeU: Visually interpretable deep learning for mortality prediction inside the ICU
- Deep Learning for Medical Image Segmentation
- Practical Black-Box Attacks against Machine Learning
- Taurus: A Data Plane Architecture for Per-Packet ML
- Iterative Machine Teaching
- The Power of Sparsity in Convolutional Neural Networks
- Cronus: Robust and Heterogeneous Collaborative Learning with Black-Box Knowledge Transfer
- Stealing Neural Networks via Timing Side Channels
- Benchmarking Deep Learning Interpretability in Time Series Predictions
- A topological insight into restricted Boltzmann machines
- Supervised Compression for Resource-Constrained Edge Computing Systems
- Interpretation of Neural Networks is Fragile
- DyNet: Dynamic Convolution for Accelerating Convolutional Neural Networks
- Giraffe: Using Deep Reinforcement Learning to Play Chess
- Domain Robustness in Neural Machine Translation
- Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding Systems
- Neural Compression and Filtering for Edge-assisted Real-time Object Detection in Challenged Networks
- ResKD: Residual-Guided Knowledge Distillation
- Ensemble Distillation for Neural Machine Translation
- Bi-GCN: Binary Graph Convolutional Network
- Lets keep it simple, Using simple architectures to outperform deeper and more complex architectures
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
- Learning From Noisy Large-Scale Datasets With Minimal Supervision
- The life of a New York City noise sensor network
- Relational Knowledge Distillation
- Deep Mutual Learning
- Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion
- 2PFPCE: Two-Phase Filter Pruning Based on Conditional Entropy
- NoScope: Optimizing Neural Network Queries over Video at Scale
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective
- On the Compression of Recurrent Neural Networks with an Application to LVCSR acoustic modeling for Embedded Speech Recognition
- Exploring Linear Relationship in Feature Map Subspace for ConvNets Compression
- PoPS: Policy Pruning and Shrinking for Deep Reinforcement Learning
- Direct Speech-to-image Translation
- Channel Distillation: Channel-Wise Attention for Knowledge Distillation
- Kickstarting Deep Reinforcement Learning
- What Kinds of Functions do Deep Neural Networks Learn? Insights from Variational Spline Theory
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- Distilling Object Detectors with Feature Richness
- Variational Information Distillation for Knowledge Transfer
- Hydra: Preserving Ensemble Diversity for Model Distillation
- Toward Fast and Accurate Neural Chinese Word Segmentation with Multi-Criteria Learning
- Predicting Deeper into the Future of Semantic Segmentation
- Zero-shot Knowledge Transfer via Adversarial Belief Matching
- Considerations When Learning Additive Explanations for Black-Box Models
- Compressing Convolutional Neural Networks
- DSConv: Efficient Convolution Operator
- Rain O'er Me: Synthesizing real rain to derain with data distillation
- Deep Learning Meets Sparse Regularization: A Signal Processing Perspective
- Transfer Learning for Speech and Language Processing
- Log-DenseNet: How to Sparsify a DenseNet
- Mix&Match - Agent Curricula for Reinforcement Learning
- ReachNN: Reachability Analysis of Neural-Network Controlled Systems
- ShrinkTeaNet: Million-scale Lightweight Face Recognition via Shrinking Teacher-Student Networks
- Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial Networks
- Blockwisely Supervised Neural Architecture Search with Knowledge Distillation
- A Large RGB-D Dataset for Semi-supervised Monocular Depth Estimation
- Addressing Missing Labels in Large-Scale Sound Event Recognition Using a Teacher-Student Framework With Loss Masking
- Optimal Subarchitecture Extraction For BERT
- FEED: Feature-level Ensemble for Knowledge Distillation
- SAGE: A Split-Architecture Methodology for Efficient End-to-End Autonomous Vehicle Control
- Embedding Deep Networks into Visual Explanations
- Active Long Term Memory Networks
- Towards Effective Low-bitwidth Convolutional Neural Networks
- Dataset Distillation
- Distilled Neural Networks for Efficient Learning to Rank
- A Closer Look at Structured Pruning for Neural Network Compression
- In Teacher We Trust: Learning Compressed Models for Pedestrian Detection
- Accelerate CNNs from Three Dimensions: A Comprehensive Pruning Framework
- Distilling Knowledge from Graph Convolutional Networks
- DBP: Discrimination Based Block-Level Pruning for Deep Model Acceleration
- Residual Knowledge Distillation
- Universal Approximation Power of Deep Residual Neural Networks via Nonlinear Control Theory
- Cracking the Black Box: Distilling Deep Sports Analytics
- Ranking to Learn and Learning to Rank: On the Role of Ranking in Pattern Recognition Applications
- Self-supervised Moving Vehicle Tracking with Stereo Sound
- Black-Box Ripper: Copying black-box models using generative evolutionary algorithms
- Joint Neural Architecture Search and Quantization
- Ranking Distillation: Learning Compact Ranking Models With High Performance for Recommender System
- Modeling Lost Information in Lossy Image Compression
- Preparing Lessons: Improve Knowledge Distillation with Better Supervision
- Deep Architectures for Modulation Recognition
- Understanding symmetries in deep networks
- Label-guided Attention Distillation for Lane Segmentation
- MKD: a Multi-Task Knowledge Distillation Approach for Pretrained Language Models
- Efficient Representation of Low-Dimensional Manifolds using Deep Networks
- Fast ConvNets Using Group-wise Brain Damage
- Optimizing speed/accuracy trade-off for person re-identification via knowledge distillation
- Copycat CNN: Are Random Non-Labeled Data Enough to Steal Knowledge from Black-box Models?
- Model Distillation with Knowledge Transfer from Face Classification to Alignment and Verification
- Why distillation helps: a statistical perspective
- Dynamic Hard Pruning of Neural Networks at the Edge of the Internet
- Structured Knowledge Distillation for Dense Prediction
- Interpreting Deep Classifier by Visual Distillation of Dark Knowledge
- Few Sample Knowledge Distillation for Efficient Network Compression
- ThUnderVolt: Enabling Aggressive Voltage Underscaling and Timing Error Resilience for Energy Efficient Deep Neural Network Accelerators
- Deep Adaptive Temporal Pooling for Activity Recognition
- RGB-based 3D Hand Pose Estimation via Privileged Learning with Depth Images
- Exact and Consistent Interpretation for Piecewise Linear Neural Networks: A Closed Form Solution
- Revisiting Knowledge Distillation via Label Smoothing Regularization
- Knowledge Distillation for End-to-End Person Search
- Distilling Word Embeddings: An Encoding Approach
- Knowledge distillation using unlabeled mismatched images
- Compact Trilinear Interaction for Visual Question Answering
- Learning Anytime Predictions in Neural Networks via Adaptive Loss Balancing
- ProSelfLC: Progressive Self Label Correction for Training Robust Deep Neural Networks
- Emotion Recognition in Speech using Cross-Modal Transfer in the Wild
- The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding
- How to Manipulate CNNs to Make Them Lie: the GradCAM Case
- A Survey on Resilient Machine Learning
- Knowledge Transfer Pre-training
- Zero-Shot Knowledge Distillation from a Decision-Based Black-Box Model
- Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
- An Experimental Study of the Impact of Pre-training on the Pruning of a Convolutional Neural Network
- Teacher-Student Compression with Generative Adversarial Networks
- Search for Better Students to Learn Distilled Knowledge
- Few-Shot Self Reminder to Overcome Catastrophic Forgetting
- Fast, Accurate, and Simple Models for Tabular Data via Augmented Distillation
- BlockSwap: Fisher-guided Block Substitution for Network Compression on a Budget
- Learning to Steer by Mimicking Features from Heterogeneous Auxiliary Networks
- Pruning via Iterative Ranking of Sensitivity Statistics
- Deep Clustered Convolutional Kernels
- CoCoPIE: Making Mobile AI Sweet As PIE --Compression-Compilation Co-Design Goes a Long Way
- Security and Privacy for Artificial Intelligence: Opportunities and Challenges
- Knowledge Squeezed Adversarial Network Compression
- Feature Representation Analysis of Deep Convolutional Neural Network using Two-stage Feature Transfer -An Application for Diffuse Lung Disease Classification-
- Discriminative Segmental Cascades for Feature-Rich Phone Recognition
- An Embarrassingly Simple Approach for Knowledge Distillation
- Modularized Morphing of Neural Networks
- Error Forward-Propagation: Reusing Feedforward Connections to Propagate Errors in Deep Learning
- Pruning artificial neural networks: a way to find well-generalizing, high-entropy sharp minima
- Efficient and Private Federated Learning with Partially Trainable Networks
- When Does Preconditioning Help or Hurt Generalization?
- Noisy Self-Knowledge Distillation for Text Summarization
- Improving the Interpretability of Deep Neural Networks with Knowledge Distillation
- Symmetry-invariant optimization in deep networks
- Balanced Knowledge Distillation for Long-tailed Learning
- Ensemble Knowledge Distillation for Learning Improved and Efficient Networks
- Variational Prototype Replays for Continual Learning
- In Defense of the Direct Perception of Affordances
- Knowledge Distillation via Weighted Ensemble of Teaching Assistants
- Share your Model instead of your Data: Privacy Preserving Mimic Learning for Ranking
- Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
- Nonparametric Neural Networks
- Kronecker Recurrent Units
- Fingerprints: Fixed Length Representation via Deep Networks and Domain Knowledge
- GASL: Guided Attention for Sparsity Learning in Deep Neural Networks
- Compact Global Descriptor for Neural Networks
- Towards Practical Lottery Ticket Hypothesis for Adversarial Training
- Distilling BERT into Simple Neural Networks with Unlabeled Transfer Data
- Interpreting Deep Neural Networks Through Variable Importance
- Hybrid Orthogonal Projection and Estimation (HOPE): A New Framework to Probe and Learn Neural Networks
- Knowledge as Priors: Cross-Modal Knowledge Generalization for Datasets without Superior Knowledge
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
- Transferring Knowledge Distillation for Multilingual Social Event Detection
- Enhancing Explainability of Neural Networks through Architecture Constraints
- Secost: Sequential co-supervision for large scale weakly labeled audio event detection
- Generative Knowledge Transfer for Neural Language Models
- Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias
- Hu-Fu: Hardware and Software Collaborative Attack Framework against Neural Networks
- Blending LSTMs into CNNs
- Privileged Knowledge Distillation for Online Action Detection
- Video Captioning with Guidance of Multimodal Latent Topics
- Collaborative Distillation for Ultra-Resolution Universal Style Transfer
- Deploy Large-Scale Deep Neural Networks in Resource Constrained IoT Devices with Local Quantization Region
- Deep ReLU Networks Preserve Expected Length
- The Shallow End: Empowering Shallower Deep-Convolutional Networks through Auxiliary Outputs
- Interactive Knowledge Distillation
- Ensemble Transformer for Efficient and Accurate Ranking Tasks: an Application to Question Answering Systems
- Model Complexity of Deep Learning: A Survey
- A Scale Mixture Perspective of Multiplicative Noise in Neural Networks
- ClusterFit: Improving Generalization of Visual Representations
- Real-time Memory Efficient Large-pose Face Alignment via Deep Evolutionary Network
- PRGFlow: Benchmarking SWAP-Aware Unified Deep Visual Inertial Odometry
- Shapley Value as Principled Metric for Structured Network Pruning
- Human-in-the-loop Extraction of Interpretable Concepts in Deep Learning Models
- ESAI: Efficient Split Artificial Intelligence via Early Exiting Using Neural Architecture Search
- Techniques for Symbol Grounding with SATNet
- Structural Knowledge Distillation: Tractably Distilling Information for Structured Predictor
- Learning From Less Data: Diversified Subset Selection and Active Learning in Image Classification Tasks
- Cross-lingual Distillation for Text Classification
- Sparsifying Neural Network Connections for Face Recognition
- On the Depth of Deep Neural Networks: A Theoretical View
- Even your Teacher Needs Guidance: Ground-Truth Targets Dampen Regularization Imposed by Self-Distillation
- FATE: Fast and Accurate Timing Error Prediction Framework for Low Power DNN Accelerator Design
- On-the-fly Network Pruning for Object Detection
- Exploiting Channel Similarity for Accelerating Deep Convolutional Neural Networks
- Data Efficient Stagewise Knowledge Distillation
- Adversarial Learning with Margin-based Triplet Embedding Regularization
- Distilling with Performance Enhanced Students
- On the Turnpike to Design of Deep Neural Nets: Explicit Depth Bounds
- Online Knowledge Distillation via Multi-branch Diversity Enhancement
- Knowledge Transfer via Dense Cross-Layer Mutual-Distillation
- Sequence Prediction with Neural Segmental Models
- Distilling Knowledge for Search-based Structured Prediction
- A Simple but Effective BERT Model for Dialog State Tracking on Resource-Limited Systems
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set Synthesis
- SAIA: Split Artificial Intelligence Architecture for Mobile Healthcare System
- ImageNet-21K Pretraining for the Masses
- Distilling Audio-Visual Knowledge by Compositional Contrastive Learning
- Distilling Knowledge Using Parallel Data for Far-field Speech Recognition
- Online Knowledge Distillation with Diverse Peers
- Cross-Layer Distillation with Semantic Calibration
- MV-MR: multi-views and multi-representations for self-supervised learning and knowledge distillation
- Model Extraction Attacks against Recurrent Neural Networks
- Leveraging Just a Few Keywords for Fine-Grained Aspect Detection Through Weakly Supervised Co-Training
- Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN
- CEKD:Cross Ensemble Knowledge Distillation for Augmented Fine-grained Data
- Machine Learning Automation Toolbox (MLaut)
- Improving the Learning of Multi-column Convolutional Neural Network for Crowd Counting
- Embedded Knowledge Distillation in Depth-Level Dynamic Neural Network
- SECS: Efficient Deep Stream Processing via Class Skew Dichotomy
- A Case for Backward Compatibility for Human-AI Teams
- Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge Distillation
- Structure-Level Knowledge Distillation For Multilingual Sequence Labeling
- Nonparametric Bayesian Deep Networks with Local Competition
- Sketching and Neural Networks
- Regularize, Expand and Compress: Multi-task based Lifelong Learning via NonExpansive AutoML
- Everything old is new again: A multi-view learning approach to learning using privileged information and distillation
- Extreme Low Resolution Activity Recognition with Confident Spatial-Temporal Attention Transfer
- Compact representations of convolutional neural networks via weight pruning and quantization
- Data-Driven Compression of Convolutional Neural Networks
- Toward Computation and Memory Efficient Neural Network Acoustic Models with Binary Weights and Activations
- Training convolutional neural networks with cheap convolutions and online distillation
- Cross-View Policy Learning for Street Navigation
- DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers
- Scalable Syntax-Aware Language Models Using Knowledge Distillation
- Leveraging Advantages of Interactive and Non-Interactive Models for Vector-Based Cross-Lingual Information Retrieval
- Robustness and Diversity Seeking Data-Free Knowledge Distillation
- A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models
- ADD: Augmented Disentanglement Distillation Framework for Improving Stock Trend Forecasting
- Accelerating Large Scale Knowledge Distillation via Dynamic Importance Sampling
- How low can you go? Privacy-preserving people detection with an omni-directional camera
- ECC: Platform-Independent Energy-Constrained Deep Neural Network Compression via a Bilinear Regression Model
- Student Network Learning via Evolutionary Knowledge Distillation
- YASENN: Explaining Neural Networks via Partitioning Activation Sequences
- Fast Human Pose Estimation
- Principal Component Networks: Parameter Reduction Early in Training
- Modality Distillation with Multiple Stream Networks for Action Recognition
- On the Orthogonality of Knowledge Distillation with Other Techniques: From an Ensemble Perspective
- Detecting Forged Facial Videos using convolutional neural network
- Paying more attention to snapshots of Iterative Pruning: Improving Model Compression via Ensemble Distillation
- Convolutional Tables Ensemble: classification in microseconds
- Oracle Teacher: Leveraging Target Information for Better Knowledge Distillation of CTC Models
- Leveraging Undiagnosed Data for Glaucoma Classification with Teacher-Student Learning
- Learning from Noisy Labels with Noise Modeling Network
- Spirit Distillation: Precise Real-time Semantic Segmentation of Road Scenes with Insufficient Data
- Deep Transfer Learning with Ridge Regression
- Incremental Meta-Learning via Indirect Discriminant Alignment
- BasisConv: A method for compressed representation and learning in CNNs
- EagleEye: Attack-Agnostic Defense against Adversarial Inputs (Technical Report)
- Revisiting the dynamics of Bose-Einstein condensates in a double well by deep learning with a hybrid network
- Developing efficient transfer learning strategies for robust scene recognition in mobile robotics using pre-trained convolutional neural networks
- The Unreasonable Effectiveness of Patches in Deep Convolutional Kernels Methods
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge Distillation
- Deep Epitome for Unravelling Generalized Hamming Network: A Fuzzy Logic Interpretation of Deep Learning
- Deep Neural Network Approximation using Tensor Sketching
- DarkGAN: Exploiting Knowledge Distillation for Comprehensible Audio Synthesis with GANs
- Supervised Robustness-preserving Data-free Neural Network Pruning
- Dynamic-TinyBERT: Boost TinyBERT's Inference Efficiency by Dynamic Sequence Length
- Lightweight 3D Human Pose Estimation Network Training Using Teacher-Student Learning
- Fixing the Teacher-Student Knowledge Discrepancy in Distillation
- Are wider nets better given the same number of parameters?
- Stochastic Model Pruning via Weight Dropping Away and Back
- Efficient detection of adversarial images
- Distilling Pixel-Wise Feature Similarities for Semantic Segmentation
- Sample Efficient Learning of Image-Based Diagnostic Classifiers Using Probabilistic Labels
- Contrastive Semi-supervised Learning for ASR
- Efficient Object Embedding for Spliced Image Retrieval
- MLitB: Machine Learning in the Browser
- Knowledge Distillation for Small-footprint Highway Networks
- Designing Interpretable Approximations to Deep Reinforcement Learning
- Learning to Augment for Data-Scarce Domain BERT Knowledge Distillation
- Blind Adversarial Pruning: Balance Accuracy, Efficiency and Robustness
- Teacher-Student Domain Adaptation for Biosensor Models
- Domain Adaptation Regularization for Spectral Pruning
- Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models' Carbon Footprint
- From Consensus to Disagreement: Multi-Teacher Distillation for Semi-Supervised Relation Extraction
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- Adaptive Distillation: Aggregating Knowledge from Multiple Paths for Efficient Distillation
- Compositionality Through Language Transmission, using Artificial Neural Networks
- Label Augmentation via Time-based Knowledge Distillation for Financial Anomaly Detection
- Similarity Transfer for Knowledge Distillation
- Improving Span-based Question Answering Systems with Coarsely Labeled Data
- Mixture of Expert/Imitator Networks: Scalable Semi-supervised Learning Framework
- Dynamic Slimmable Network
- A New Training Framework for Deep Neural Network
- Advancing Multi-Accented LSTM-CTC Speech Recognition using a Domain Specific Student-Teacher Learning Paradigm
- RDPD: Rich Data Helps Poor Data via Imitation
- Learning in Deep Neural Networks Using a Biologically Inspired Optimizer
- Robust Student Network Learning
- A Framework for Behavioral Biometric Authentication using Deep Metric Learning on Mobile Devices
- Prime-Aware Adaptive Distillation
- SoFAr: Shortcut-based Fractal Architectures for Binary Convolutional Neural Networks
- Sequence Training and Adaptation of Highway Deep Neural Networks
- Progressive Label Distillation: Learning Input-Efficient Deep Neural Networks
- Benchmarking Approximate Inference Methods for Neural Structured Prediction
- The Pupil Has Become the Master: Teacher-Student Model-Based Word Embedding Distillation with Ensemble Learning
- Knowledge Distillation-aided End-to-End Learning for Linear Precoding in Multiuser MIMO Downlink Systems with Finite-Rate Feedback
- A Fully Tensorized Recurrent Neural Network
- Representation Evaluation Block-based Teacher-Student Network for the Industrial Quality-relevant Performance Modeling and Monitoring
- On Estimating the Training Cost of Conversational Recommendation Systems
- Robust testing of low-dimensional functions
- Transferring Inter-Class Correlation
- Towards Modality Transferable Visual Information Representation with Optimal Model Compression
- Sketching Linear Classifiers over Data Streams
- Channel Planting for Deep Neural Networks using Knowledge Distillation
- FastSal: a Computationally Efficient Network for Visual Saliency Prediction
- SmartDeal: Re-Modeling Deep Network Weights for Efficient Inference and Training
- Multilingual AMR Parsing with Noisy Knowledge Distillation
- Compressing Deep Convolutional Neural Networks by Stacking Low-dimensional Binary Convolution Filters
- Low-Latency Incremental Text-to-Speech Synthesis with Distilled Context Prediction Network
- A Studious Approach to Semi-Supervised Learning
- Interpretable Few-Shot Learning via Linear Distillation
- Secure Your Ride: Real-time Matching Success Rate Prediction for Passenger-Driver Pairs
- Inner Ensemble Networks: Average Ensemble as an Effective Regularizer
- Exact and Consistent Interpretation of Piecewise Linear Models Hidden behind APIs: A Closed Form Solution
- Implicit Priors for Knowledge Sharing in Bayesian Neural Networks
- AIP: Adversarial Iterative Pruning Based on Knowledge Transfer for Convolutional Neural Networks
- Learning Energy-Based Approximate Inference Networks for Structured Applications in NLP
- Justlookup: One Millisecond Deep Feature Extraction for Point Clouds By Lookup Tables
- Long Short-Term Sample Distillation
- Towards Real-time Mispronunciation Detection in Kids' Speech
- Learning Fast Matching Models from Weak Annotations
- Domain Adaptation for Facial Expression Classifier via Domain Discrimination and Gradient Reversal
- Person Identification Based on Hand Tremor Characteristics
- Inference of a Multi-Domain Machine Learning Model to Predict Mortality in Hospital Stays for Patients with Cancer upon Febrile Neutropenia Onset
- Reducing the Deployment-Time Inference Control Costs of Deep Reinforcement Learning Agents via an Asymmetric Architecture
- RGP: Neural Network Pruning through Its Regular Graph Structure
- Meta-Teacher For Face Anti-Spoofing
- A Survey on Green Deep Learning
- Be Your Own Best Competitor! Multi-Branched Adversarial Knowledge Transfer
- A Note on Knowledge Distillation Loss Function for Object Classification
- Shakeout: A New Approach to Regularized Deep Neural Network Training
- On the Efficiency of Subclass Knowledge Distillation in Classification Tasks
- Collaborative Group Learning
- Is the Meta-Learning Idea Able to Improve the Generalization of Deep Neural Networks on the Standard Supervised Learning?
- Perceptual Gradient Networks
- Network Implosion: Effective Model Compression for ResNets via Static Layer Pruning and Retraining
- A copula-based visualization technique for a neural network
- Towards glass-box CNNs
- ADA-Tucker: Compressing Deep Neural Networks via Adaptive Dimension Adjustment Tucker Decomposition
- A One-step Pruning-recovery Framework for Acceleration of Convolutional Neural Networks
- Teach an all-rounder with experts in different domains
- Hardware-Software Codesign of Accurate, Multiplier-free Deep Neural Networks
- Triplet Loss for Knowledge Distillation
- CHEER: Rich Model Helps Poor Model via Knowledge Infusion
- Feature Statistics Guided Efficient Filter Pruning
- Matrix Factorization on GPUs with Memory Optimization and Approximate Computing
- Knowledge Distillation By Sparse Representation Matching
- Dynamic Domain Adaptation for Efficient Inference
- Self-Referenced Deep Learning
- Self-Teaching Machines to Read and Comprehend with Large-Scale Multi-Subject Question-Answering Data