CLIP-Art: Contrastive Pre-training for Fine-Grained Art Classification
arXiv:2204.14244 · doi:10.1109/CVPRW53098.2021.00444
Abstract
Existing computer vision research in artwork struggles with artwork's fine-grained attributes recognition and lack of curated annotated datasets due to their costly creation. To the best of our knowledge, we are one of the first methods to use CLIP (Contrastive Language-Image Pre-Training) to train a neural network on a variety of artwork images and text descriptions pairs. CLIP is able to learn directly from free-form art descriptions, or, if available, curated fine-grained labels. Model's zero-shot capability allows predicting accurate natural language description for a given image, without directly optimizing for the task. Our approach aims to solve 2 challenges: instance retrieval and fine-grained artwork attribute recognition. We use the iMet Dataset, which we consider the largest annotated artwork dataset. In this benchmark we achieved competitive results using only self-supervision.
CVPR CVFAD Workshop 2021
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A Simple Framework for Contrastive Learning of Visual Representations
- Representation Learning with Contrastive Predictive Coding
- Learning Transferable Visual Models From Natural Language Supervision
- Deep Residual Learning for Image Recognition
- Fine-Grained Visual Classification of Aircraft
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- On the Variance of the Adaptive Learning Rate and Beyond
- Lookahead Optimizer: k steps forward, 1 step back
- See Better Before Looking Closer: Weakly Supervised Data Augmentation Network for Fine-Grained Visual Classification
- Fine-grained Visual-textual Representation Learning
- The iMet Collection 2019 Challenge Dataset
Cited by in corpus (7)
- GalleryGPT: Analyzing Paintings with Large Multimodal Models
- Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
- Multimodal Metadata Assignment for Cultural Heritage Artifacts
- LM-CPPF: Paraphrasing-Guided Data Augmentation for Contrastive Prompt-Based Few-Shot Fine-Tuning
- Toward the Automated Localization of Buggy Mobile App UIs from Bug Descriptions
- SOOD-ImageNet: a Large-Scale Dataset for Semantic Out-Of-Distribution Image Classification and Semantic Segmentation
- An implementation of the "Guess who?" game using CLIP