Reproducible scaling laws for contrastive language-image learning
arXiv:2212.07143 · doi:10.1109/CVPR52729.2023.00276
Abstract
Scaling up neural networks has led to remarkable performance across a wide range of tasks. Moreover, performance often follows reliable scaling laws as a function of training set size, model size, and compute, which offers valuable guidance as large-scale experiments are becoming increasingly expensive. However, previous work on scaling laws has primarily used private data \& models or focused on uni-modal language or vision learning. To address these limitations, we investigate scaling laws for contrastive language-image pre-training (CLIP) with the public LAION dataset and the open-source OpenCLIP repository. Our large-scale experiments involve models trained on up to two billion image-text pairs and identify power law scaling for multiple downstream tasks including zero-shot classification, retrieval, linear probing, and end-to-end fine-tuning. We find that the training distribution plays a key role in scaling laws as the OpenAI and OpenCLIP models exhibit different scaling behavior despite identical model architectures and similar training recipes. We open-source our evaluation workflow and all models, including the largest public CLIP models, to ensure reproducibility and make scaling laws research more accessible. Source code and instructions to reproduce this study will be available at https://github.com/LAION-AI/scaling-laws-openclip
CVPR 2023. Version with minor extension. Original: https://openaccess.thecvf.com/content/CVPR2023/html/Cherti_Reproducible_Scaling_Laws_for_Contrastive_Language-Image_Learning_CVPR_2023_paper
References in corpus (32)
- Adam: A Method for Stochastic Optimization
- Decoupled Weight Decay Regularization
- Representation Learning with Contrastive Predictive Coding
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- YFCC100M: The New Data in Multimedia Research
- Reconciling modern machine learning practice and the bias-variance trade-off
- Flamingo: a Visual Language Model for Few-Shot Learning
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Robust Speech Recognition via Large-Scale Weak Supervision
- LAION-5B: An open large-scale dataset for training next generation image-text models
- BEiT: BERT Pre-Training of Image Transformers
- Training Compute-Optimal Large Language Models
- CoCa: Contrastive Captioners are Image-Text Foundation Models
- Deep Learning Scaling is Predictable, Empirically
- Do ImageNet Classifiers Generalize to ImageNet?
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Imagen Video: High Definition Video Generation with Diffusion Models
- Learning Robust Global Representations by Penalizing Local Predictive Power
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- Measuring Robustness to Natural Distribution Shifts in Image Classification
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Scaling Laws for Autoregressive Generative Modeling
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- Beyond neural scaling laws: beating power law scaling via data pruning
- Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts
- Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
- Accuracy on the Line: On the Strong Correlation Between Out-of-Distribution and In-Distribution Generalization
- Patching open-vocabulary models by interpolating weights
- Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)
- Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP
- On the Predictability of Pruning Across Scales
- Where Should I Spend My FLOPS? Efficiency Evaluations of Visual Pre-training Methods
Cited by in corpus (29)
- DINOv2: Learning Robust Visual Features without Supervision
- A Survey on Multimodal Large Language Models
- CLIP in Medical Imaging: A Survey
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Learning without Forgetting for Vision-Language Models
- Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset
- RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
- Does CLIP Know My Face?
- Who's in and who's out? A case study of multimodal CLIP-filtering in DataComp
- Application-Driven Exascale: The JUPITER Benchmark Suite
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
- Multilingual Vision-Language Pre-training for the Remote Sensing Domain
- Towards Real Zero-Shot Camouflaged Object Segmentation without Camouflaged Annotations
- Evaluating the Fairness of Discriminative Foundation Models in Computer Vision
- Quadratic Neuron-empowered Heterogeneous Autoencoder for Unsupervised Anomaly Detection
- Climber: Toward Efficient Scaling Laws for Large Recommendation Models
- Parameter-aware high-fidelity microstructure generation using stable diffusion
- Multi-label Cluster Discrimination for Visual Representation Learning
- The Impacts of Data, Ordering, and Intrinsic Dimensionality on Recall in Hierarchical Navigable Small Worlds
- Zero-shot detection of buildings in mobile LiDAR using Language Vision Model
- Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
- Human-in-the-loop Reasoning For Traffic Sign Detection: Collaborative Approach Yolo With Video-llava
- Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
- OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
- Leveraging Self-Supervised Vision Transformers for Segmentation-based Transfer Function Design
- HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion
- Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance
- Training-free Temporal Object Tracking in Surgical Videos
- A Matter of Time: Revealing the Structure of Time in Vision-Language Models