Deeply Learning the Messages in Message Passing Inference
arXiv:1506.02108
Abstract
Deep structured output learning shows great promise in tasks like semantic image segmentation. We proffer a new, efficient deep structured model learning scheme, in which we show how deep Convolutional Neural Networks (CNNs) can be used to estimate the messages in message passing inference for structured prediction with Conditional Random Fields (CRFs). With such CNN message estimators, we obviate the need to learn or evaluate potential functions for message calculation. This confers significant efficiency for learning, since otherwise when performing structured learning for a CRF with CNN potentials it is necessary to undertake expensive inference for every stochastic gradient iteration. The network output dimension for message estimation is the same as the number of classes, in contrast to the network output for general CNN potential functions in CRFs, which is exponential in the order of the potentials. Hence CNN message learning has fewer network parameters and is more scalable for cases that a large number of classes are involved. We apply our method to semantic image segmentation on the PASCAL VOC 2012 dataset. We achieve an intersection-over-union score of 73.4 on its test set, which is the best reported result for methods using the VOC training images alone. This impressive performance demonstrates the effectiveness and usefulness of our CNN message learning method.
11 pages. Appearing in Proc. The Twenty-ninth Annual Conference on Neural Information Processing Systems (NIPS), 2015, Montreal, Canada
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Conditional Random Fields as Recurrent Neural Networks
- Learning Depth from Single Monocular Images Using Deep Convolutional Neural Fields
- Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation
- Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation
- Fully Connected Deep Structured Networks
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Piecewise Training for Undirected Models
- Efficient piecewise training of deep structured models for semantic segmentation
- Feedforward semantic segmentation with zoom-out features
Cited by in corpus (10)
- Discriminative Embeddings of Latent Variable Models for Structured Data
- Structured prediction models for RNN based sequence labeling in clinical text
- Reasoning Visual Dialogs with Structural and Partial Observations
- Discriminative Training of Deep Fully-connected Continuous CRF with Task-specific Loss
- Learning to Filter with Predictive State Inference Machines
- PDP: A General Neural Framework for Learning Constraint Satisfaction Solvers
- Amortized Bethe Free Energy Minimization for Learning MRFs
- Neuralizing Efficient Higher-order Belief Propagation
- Aerial Images Meet Crowdsourced Trajectories: A New Approach to Robust Road Extraction
- Belief Propagation for Approximate Inference