LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
arXiv:2111.02114
Abstract
Multi-modal language-vision models trained on hundreds of millions of image-text pairs (e.g. CLIP, DALL-E) gained a recent surge, showing remarkable capability to perform zero- or few-shot learning and transfer even in absence of per-sample labels on target image data. Despite this trend, to date there has been no publicly available datasets of sufficient scale for training such models from scratch. To address this issue, in a community effort we build and release for public LAION-400M, a dataset with CLIP-filtered 400 million image-text pairs, their CLIP embeddings and kNN indices that allow efficient similarity search.
Short version. Accepted at Data Centric AI NeurIPS Workshop 2021
References in corpus (8)
- Learning Transferable Visual Models From Natural Language Supervision
- Language Models are Few-Shot Learners
- Scaling Laws for Neural Language Models
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Zero-Shot Text-to-Image Generation
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Scaling Laws for Autoregressive Generative Modeling
- One Epoch Is All You Need
Cited by in corpus (26)
- DINOv2: Learning Robust Visual Features without Supervision
- Reproducible scaling laws for contrastive language-image learning
- AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection
- DiffEdit: Diffusion-based semantic image editing with mask guidance
- Data Governance in the Age of Large-Scale Data-Driven Language Technology
- Vector Quantized Diffusion Model for Text-to-Image Synthesis
- Hallucination Detection in Foundation Models for Decision-Making: A Flexible Definition and Review of the State of the Art
- Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- Several categories of Large Language Models (LLMs): A Short Survey
- Re-Scoring Using Image-Language Similarity for Few-Shot Object Detection
- Scaling Up Vision-Language Pre-training for Image Captioning
- Does CLIP Know My Face?
- RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
- Situating the social issues of image generation models in the model life cycle: a sociotechnical approach
- Evaluating the Fairness of Discriminative Foundation Models in Computer Vision
- Image-Text Pre-Training for Logo Recognition
- Multimedia Generative Script Learning for Task Planning
- fruit-SALAD: A Style Aligned Artwork Dataset to reveal similarity perception in image embeddings
- Toward High Quality Facial Representation Learning
- Face Aging via Diffusion-based Editing
- Centered Masking for Language-Image Pre-Training
- The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain Adaptation
- Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval
- Weighted Ensemble Models Are Strong Continual Learners
- GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models