Visual Representations: Defining Properties and Deep Approximations
arXiv:1411.7676
Abstract
Visual representations are defined in terms of minimal sufficient statistics of visual data, for a class of tasks, that are also invariant to nuisance variability. Minimal sufficiency guarantees that we can store a representation in lieu of raw data with smallest complexity and no performance loss on the task at hand. Invariance guarantees that the statistic is constant with respect to uninformative transformations of the data. We derive analytical expressions for such representations and show they are related to feature descriptors commonly used in computer vision, as well as to convolutional neural networks. This link highlights the assumptions and approximations tacitly assumed by these methods and explains empirical practices such as clamping, pooling and joint normalization.
UCLA CSD TR140023, Nov. 12, 2014, revised April 13, 2015, November 13, 2015, February 28, 2016
References in corpus (9)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Descriptor Matching with Convolutional Neural Networks: a Comparison to SIFT
- Understanding Deep Image Representations by Inverting Them
- A Probabilistic Theory of Deep Learning
- Deformable Part Models are Convolutional Neural Networks
- Sketch-a-Net that Beats Humans
- On Invariance and Selectivity in Representation Learning
- Multi-scale Orderless Pooling of Deep Convolutional Activation Features
- Domain-Size Pooling in Local Descriptors: DSP-SIFT
Cited by in corpus (6)
- Invariant Representations without Adversarial Training
- Tiling and Stitching Segmentation Output for Remote Sensing: Basic Challenges and Recommendations
- Translation Insensitive CNNs
- Evolution Is All You Need: Phylogenetic Augmentation for Contrastive Learning
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation
- A Cyclically-Trained Adversarial Network for Invariant Representation Learning