Looking at Outfit to Parse Clothing
arXiv:1703.01386
Abstract
This paper extends fully-convolutional neural networks (FCN) for the clothing parsing problem. Clothing parsing requires higher-level knowledge on clothing semantics and contextual cues to disambiguate fine-grained categories. We extend FCN architecture with a side-branch network which we refer outfit encoder to predict a consistent set of clothing labels to encourage combinatorial preference, and with conditional random field (CRF) to explicitly consider coherent label assignment to the given image. The empirical results using Fashionista and CFPD datasets show that our model achieves state-of-the-art performance in clothing parsing, without additional supervision during training. We also study the qualitative influence of annotation on the current clothing parsing benchmarks, with our Web-based tool for multi-scale pixel-wise annotation and manual refinement effort to the Fashionista dataset. Finally, we show that the image representation of the outfit encoder is useful for dress-up image retrieval application.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Learning Deconvolution Network for Semantic Segmentation
- PixelNet: Towards a General Pixel-level Architecture
Cited by in corpus (6)
- A Deep-Learning-Based Fashion Attributes Detection Model
- fAshIon after fashion: A Report of AI in Fashion
- Pose Guided Fashion Image Synthesis Using Deep Generative Model
- Searching for Apparel Products from Images in the Wild
- Recommending Outfits from Personal Closet
- Recommendation or Discrimination?: Quantifying Distribution Parity in Information Retrieval Systems