Automatic Concept Discovery from Parallel Text and Visual Corpora
arXiv:1509.07225
Abstract
Humans connect language and vision to perceive the world. How to build a similar connection for computers? One possible way is via visual concepts, which are text terms that relate to visually discriminative entities. We propose an automatic visual concept discovery algorithm using parallel text and visual corpora; it filters text terms based on the visual discriminative power of the associated images, and groups them into concepts using visual and semantic similarities. We illustrate the applications of the discovered concepts using bidirectional image and sentence retrieval task and image tagging task, and show that the discovered concepts not only outperform several large sets of manually selected concepts significantly, but also achieves the state-of-the-art performance in the retrieval task.
To appear in ICCV 2015
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Rich feature hierarchies for accurate object detection and semantic segmentation
- Explain Images with Multimodal Recurrent Neural Networks
- Learning a Recurrent Visual Representation for Image Caption Generation
- ConceptLearner: Discovering Visual Concepts from Weakly Labeled Image Collections
Cited by in corpus (10)
- ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering
- Automatic Spatially-aware Fashion Concept Discovery
- Learning Attributes Equals Multi-Source Domain Generalization
- VQS: Linking Segmentations to Questions and Answers for Supervised Attention in VQA and Question-Focused Semantic Segmentation
- A Comprehensive Survey of Deep Learning for Image Captioning
- Unsupervised Category Discovery via Looped Deep Pseudo-Task Optimization Using a Large Scale Radiology Image Database
- Learning without Prejudice: Avoiding Bias in Webly-Supervised Action Recognition
- Complex Event Recognition from Images with Few Training Examples
- Learning Action Concept Trees and Semantic Alignment Networks from Image-Description Data
- Automatic Visual Theme Discovery from Joint Image and Text Corpora