Stacked What-Where Auto-encoders
arXiv:1506.02351
Abstract
We present a novel architecture, the "stacked what-where auto-encoders" (SWWAE), which integrates discriminative and generative pathways and provides a unified approach to supervised, semi-supervised and unsupervised learning without relying on sampling during training. An instantiation of SWWAE uses a convolutional net (Convnet) (LeCun et al. (1998)) to encode the input, and employs a deconvolutional net (Deconvnet) (Zeiler et al. (2010)) to produce the reconstruction. The objective function includes reconstruction terms that induce the hidden states in the Deconvnet to be similar to those of the Convnet. Each pooling layer produces two sets of variables: the "what" which are fed to the next layer, and its complementary variable "where" that are fed to the corresponding layer in the generative decoder.
Workshop track - ICLR 2016
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Improving neural networks by preventing co-adaptation of feature detectors
- Striving for Simplicity: The All Convolutional Net
- Training Very Deep Networks
- Fractional Max-Pooling
- Fast Inference in Sparse Coding Algorithms with Applications to Object Recognition
- Winner-Take-All Autoencoders
- Lateral Connections in Denoising Autoencoders Support Supervised Learning
Cited by in corpus (69)
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Energy-based Generative Adversarial Network
- Adversarially Learned Inference
- Autoencoding beyond pixels using a learned similarity metric
- MixMatch: A Holistic Approach to Semi-Supervised Learning
- Semi-Supervised Learning with Ladder Networks
- Stacked Hourglass Networks for Human Pose Estimation
- Neural Photo Editing with Introspective Adversarial Networks
- A selectional auto-encoder approach for document image binarization
- ReMixMatch: Semi-Supervised Learning with Distribution Alignment and Augmentation Anchoring
- Understanding and Improving Convolutional Neural Networks via Concatenated Rectified Linear Units
- Loss-Sensitive Generative Adversarial Networks on Lipschitz Densities
- Semi-supervised Learning with GANs: Manifold Invariance with Improved Inference
- Augmenting Supervised Neural Networks with Unsupervised Objectives for Large-scale Image Classification
- Unsupervised Embedding Learning via Invariant and Spreading Instance Feature
- Scattering Networks for Hybrid Representation Learning
- Deep generative-contrastive networks for facial expression recognition
- Convolutional Clustering for Unsupervised Learning
- Analysis and Optimization of Convolutional Neural Network Architectures
- Deep Clustering via Joint Convolutional Autoencoder Embedding and Relative Entropy Minimization
- Learning by Association - A versatile semi-supervised training method for neural networks
- A Closed-form Solution to Photorealistic Image Stylization
- ExtremeWeather: A large-scale climate dataset for semi-supervised detection, localization, and understanding of extreme weather events
- Stacked Generative Adversarial Networks
- Deep Discriminative Clustering Analysis
- Whitening for Self-Supervised Representation Learning
- Identifying and Categorizing Anomalies in Retinal Imaging Data
- Semi-supervised deep learning by metric embedding
- Deep Linear Discriminant Analysis
- Robust Federated Learning Through Representation Matching and Adaptive Hyper-parameters
- Walk and Learn: Facial Attribute Representation Learning from Egocentric Video and Contextual Data
- Revealing Fundamental Physics from the Daya Bay Neutrino Experiment using Deep Neural Networks
- Scaling the Scattering Transform: Deep Hybrid Networks
- Self-supervised learning of visual features through embedding images into text topic spaces
- EnAET: A Self-Trained framework for Semi-Supervised and Supervised Learning with Ensemble Transformations
- Deep unsupervised learning through spatial contrasting
- A Comparative Review of Recent Few-Shot Object Detection Algorithms
- Ablation of a Robot's Brain: Neural Networks Under a Knife
- Deep Co-Space: Sample Mining Across Feature Transformation for Semi-Supervised Learning
- GATCluster: Self-Supervised Gaussian-Attention Network for Image Clustering
- On the Road with 16 Neurons: Mental Imagery with Bio-inspired Deep Neural Networks
- Deep TEN: Texture Encoding Network
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- Learning the Precise Feature for Cluster Assignment
- Reinforcement Learning-Based Coverage Path Planning with Implicit Cellular Decomposition
- Faster Convergence in Deep-Predictive-Coding Networks to Learn Deeper Representations
- Manifold Adversarial Learning
- Semi-Supervised Learning with the Deep Rendering Mixture Model
- Flow Contrastive Estimation of Energy-Based Models
- RankingMatch: Delving into Semi-Supervised Learning with Consistency Regularization and Ranking Loss
- WSAM: Visual Explanations from Style Augmentation as Adversarial Attacker and Their Influence in Image Classification
- Improved Deep Learning of Object Category using Pose Information
- CompressNet: Generative Compression at Extremely Low Bitrates
- Universum Prescription: Regularization using Unlabeled Data
- Self-Supervised Visual Representations for Cross-Modal Retrieval
- S3Pool: Pooling with Stochastic Spatial Sampling
- Label Prediction Framework for Semi-Supervised Cross-Modal Retrieval
- Supervised Deep Sparse Coding Networks
- Multi-level Feature Learning on Embedding Layer of Convolutional Autoencoders and Deep Inverse Feature Learning for Image Clustering
- Soft Autoencoder and Its Wavelet Adaptation Interpretation
- Towards Robust Pattern Recognition: A Review
- Spatial contrasting for deep unsupervised learning
- ReRankMatch: Semi-Supervised Learning with Semantics-Oriented Similarity Representation
- Discriminate-and-Rectify Encoders: Learning from Image Transformation Sets
- Image Embedded Segmentation: Uniting Supervised and Unsupervised Objectives for Segmenting Histopathological Images
- Pose Augmentation: Class-agnostic Object Pose Transformation for Object Recognition
- Adversarial Ladder Networks
- A backward pass through a CNN using a generative model of its activations
- Object Parsing in Sequences Using CoordConv Gated Recurrent Networks