Reproducible scaling laws for contrastive language-image learning
arXiv:2212.07143 · doi:10.1109/CVPR52729.2023.00276
Abstract
Scaling up neural networks has led to remarkable performance across a wide range of tasks. Moreover, performance often follows reliable scaling laws as a function of training set size, model size, and compute, which offers valuable guidance as large-scale experiments are becoming increasingly expensive. However, previous work on scaling laws has primarily used private data \& models or focused on uni-modal language or vision learning. To address these limitations, we investigate scaling laws for contrastive language-image pre-training (CLIP) with the public LAION dataset and the open-source OpenCLIP repository. Our large-scale experiments involve models trained on up to two billion image-text pairs and identify power law scaling for multiple downstream tasks including zero-shot classification, retrieval, linear probing, and end-to-end fine-tuning. We find that the training distribution plays a key role in scaling laws as the OpenAI and OpenCLIP models exhibit different scaling behavior despite identical model architectures and similar training recipes. We open-source our evaluation workflow and all models, including the largest public CLIP models, to ensure reproducibility and make scaling laws research more accessible. Source code and instructions to reproduce this study will be available at https://github.com/LAION-AI/scaling-laws-openclip
CVPR 2023. Version with minor extension. Original: https://openaccess.thecvf.com/content/CVPR2023/html/Cherti_Reproducible_Scaling_Laws_for_Contrastive_Language-Image_Learning_CVPR_2023_paper
References in corpus (14)
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Robust Speech Recognition via Large-Scale Weak Supervision
- LAION-5B: An open large-scale dataset for training next generation image-text models
- Deep Learning Scaling is Predictable, Empirically
- Do ImageNet Classifiers Generalize to ImageNet?
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Imagen Video: High Definition Video Generation with Diffusion Models
- Measuring Robustness to Natural Distribution Shifts in Image Classification
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Scaling Laws for Autoregressive Generative Modeling
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- Accuracy on the Line: On the Strong Correlation Between Out-of-Distribution and In-Distribution Generalization
- Patching open-vocabulary models by interpolating weights
- Where Should I Spend My FLOPS? Efficiency Evaluations of Visual Pre-training Methods
Cited by in corpus (27)
- DINOv2: Learning Robust Visual Features without Supervision
- A Survey on Multimodal Large Language Models
- CLIP in Medical Imaging: A Survey
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Learning without Forgetting for Vision-Language Models
- Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset
- RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
- Application-Driven Exascale: The JUPITER Benchmark Suite
- Who's in and who's out? A case study of multimodal CLIP-filtering in DataComp
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
- Multilingual Vision-Language Pre-training for the Remote Sensing Domain
- Towards Real Zero-Shot Camouflaged Object Segmentation without Camouflaged Annotations
- Evaluating the Fairness of Discriminative Foundation Models in Computer Vision
- Climber: Toward Efficient Scaling Laws for Large Recommendation Models
- Parameter-aware high-fidelity microstructure generation using stable diffusion
- Multi-label Cluster Discrimination for Visual Representation Learning
- The Impacts of Data, Ordering, and Intrinsic Dimensionality on Recall in Hierarchical Navigable Small Worlds
- Zero-shot detection of buildings in mobile LiDAR using Language Vision Model
- Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
- Human-in-the-loop Reasoning For Traffic Sign Detection: Collaborative Approach Yolo With Video-llava
- Leveraging Self-Supervised Vision Transformers for Segmentation-based Transfer Function Design
- OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
- Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking
- Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance
- HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion
- A Matter of Time: Revealing the Structure of Time in Vision-Language Models
- Training-free Temporal Object Tracking in Surgical Videos