BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
arXiv:2301.12597
Abstract
The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models. This paper proposes BLIP-2, a generic and efficient pre-training strategy that bootstraps vision-language pre-training from off-the-shelf frozen pre-trained image encoders and frozen large language models. BLIP-2 bridges the modality gap with a lightweight Querying Transformer, which is pre-trained in two stages. The first stage bootstraps vision-language representation learning from a frozen image encoder. The second stage bootstraps vision-to-language generative learning from a frozen language model. BLIP-2 achieves state-of-the-art performance on various vision-language tasks, despite having significantly fewer trainable parameters than existing methods. For example, our model outperforms Flamingo80B by 8.7% on zero-shot VQAv2 with 54x fewer trainable parameters. We also demonstrate the model's emerging capabilities of zero-shot image-to-text generation that can follow natural language instructions.
Cited by in corpus (45)
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
- Woodpecker: Hallucination Correction for Multimodal Large Language Models
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text
- CFN-ESA: A Cross-Modal Fusion Network with Emotion-Shift Awareness for Dialogue Emotion Recognition
- ChatCAD+: Towards a Universal and Reliable Interactive CAD using LLMs
- Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
- Generative AI and Process Systems Engineering: The Next Frontier
- Correctness Comparison of ChatGPT-4, Gemini, Claude-3, and Copilot for Spatial Tasks
- Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
- IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models
- SoccerNet 2023 Challenges Results
- Building Privacy-Preserving and Secure Geospatial Artificial Intelligence Foundation Models
- ExpressEdit: Video Editing with Natural Language and Sketching
- Human I/O: Towards a Unified Approach to Detecting Situational Impairments
- SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
- LLaVA-Docent: Instruction Tuning with Multimodal Large Language Model to Support Art Appreciation Education
- FabGPT: An Efficient Large Multimodal Model for Complex Wafer Defect Knowledge Queries
- Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields
- Foundation Models and Transformers for Anomaly Detection: A Survey
- Automatic Medical Report Generation: Methods and Applications
- Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
- Multimodal Neural Databases
- MemoVis: A GenAI-Powered Tool for Creating Companion Reference Images for 3D Design Feedback
- PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
- Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
- LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
- Generative AI-based Prompt Evolution Engineering Design Optimization With Vision-Language Model
- UVOSAM: A Mask-free Paradigm for Unsupervised Video Object Segmentation via Segment Anything Model
- FDM-Bench: A Comprehensive Benchmark for Evaluating Large Language Models in Additive Manufacturing Tasks
- MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
- Automated Description Generation of Cytologic Findings for Lung Cytological Images Using a Pretrained Vision Model and Dual Text Decoders: Preliminary Study
- Unified Multi-modal Diagnostic Framework with Reconstruction Pre-training and Heterogeneity-combat Tuning
- PEAR: Phrase-Based Hand-Object Interaction Anticipation
- Deep Insights into Cognitive Decline: A Survey of Leveraging Non-Intrusive Modalities with Deep Learning Techniques
- Unsupervised Text Embedding Space Generation Using Generative Adversarial Networks for Text Synthesis
- HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs
- Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
- Bridging Text and Crystal Structures: Literature-driven Contrastive Learning for Materials Science
- MuG: A Multimodal Classification Benchmark on Game Data with Tabular, Textual, and Visual Fields
- AICAttack: Adversarial Image Captioning Attack with Attention-Based Optimization
- MATK: The Meme Analytical Tool Kit
- Image Captions are Natural Prompts for Text-to-Image Models
- Semantic Generative Augmentations for Few-Shot Counting