LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
arXiv:2111.02114
Abstract
Multi-modal language-vision models trained on hundreds of millions of image-text pairs (e.g. CLIP, DALL-E) gained a recent surge, showing remarkable capability to perform zero- or few-shot learning and transfer even in absence of per-sample labels on target image data. Despite this trend, to date there has been no publicly available datasets of sufficient scale for training such models from scratch. To address this issue, in a community effort we build and release for public LAION-400M, a dataset with CLIP-filtered 400 million image-text pairs, their CLIP embeddings and kNN indices that allow efficient similarity search.
Short version. Accepted at Data Centric AI NeurIPS Workshop 2021
References in corpus (7)
- Learning Transferable Visual Models From Natural Language Supervision
- Language Models are Few-Shot Learners
- Scaling Laws for Neural Language Models
- Zero-Shot Text-to-Image Generation
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Scaling Laws for Autoregressive Generative Modeling
- One Epoch Is All You Need
Cited by in corpus (6)
- DiffEdit: Diffusion-based semantic image editing with mask guidance
- Evaluating the Fairness of Discriminative Foundation Models in Computer Vision
- Image-Text Pre-Training for Logo Recognition
- Toward High Quality Facial Representation Learning
- Face Aging via Diffusion-based Editing
- The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain Adaptation