Zero-Shot Learning with Knowledge Enhanced Visual Semantic Embeddings
arXiv:2011.10889
Abstract
We improve zero-shot learning (ZSL) by incorporating common-sense knowledge in DNNs. We propose Common-Sense based Neuro-Symbolic Loss (CSNL) that formulates prior knowledge as novel neuro-symbolic loss functions that regularize visual-semantic embedding. CSNL forces visual features in the VSE to obey common-sense rules relating to hypernyms and attributes. We introduce two key novelties for improved learning: (1) enforcement of rules for a group instead of a single concept to take into account class-wise relationships, and (2) confidence margins inside logical operators that enable implicit curriculum learning and prevent premature overfitting. We evaluate the advantages of incorporating each knowledge source and show consistent gains over prior state-of-art methods in both conventional and generalized ZSL e.g. 11.5%, 5.5%, and 11.6% improvements on AWA2, CUB, and Kinetics respectively.
References in corpus (8)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Distributed Representations of Sentences and Documents
- The Kinetics Human Action Video Dataset
- Semi-Supervised Learning with Deep Generative Models
- Logic Tensor Networks: Deep Learning and Logical Reasoning from Data and Knowledge
- Leveraging the Invariant Side of Generative Zero-Shot Learning
- Visually Aligned Word Embeddings for Improving Zero-shot Learning
- Learning Structured Semantic Embeddings for Visual Recognition