Flamingo: a Visual Language Model for Few-Shot Learning
arXiv:2204.14198
Abstract
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer; captioning tasks, which evaluate the ability to describe a scene or an event; and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data.
54 pages. In Proceedings of Neural Information Processing Systems (NeurIPS) 2022
Cited by in corpus (61)
- Reproducible scaling laws for contrastive language-image learning
- Compute Trends Across Three Eras of Machine Learning
- Auditing large language models: a three-layered approach
- Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
- Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education
- OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
- Harms from Increasingly Agentic Algorithmic Systems
- Advances of Machine Learning in Materials Science: Ideas and Techniques
- A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text
- CLIP in Medical Imaging: A Survey
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- ChatCAD+: Towards a Universal and Reliable Interactive CAD using LLMs
- Exploring scalable medical image encoders beyond text supervision
- The illusion of artificial inclusion
- SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
- CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
- Deceptive AI Ecosystems: The Case of ChatGPT
- Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation
- A social path to human-like artificial intelligence
- The Survey on Multi-Source Data Fusion in Cyber-Physical-Social Systems:Foundational Infrastructure for Industrial Metaverses and Industries 5.0
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- Several categories of Large Language Models (LLMs): A Short Survey
- Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations
- SoccerNet 2023 Challenges Results
- Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models
- Exploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning
- A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis
- Does CLIP Know My Face?
- Foundation Models and Transformers for Anomaly Detection: A Survey
- Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
- RTQ: Rethinking Video-language Understanding Based on Image-text Model
- Power-LLaVA: Large Language and Vision Assistant for Power Transmission Line Inspection
- Large-vocabulary forensic pathological analyses via prototypical cross-modal contrastive learning
- Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D Representation
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real World
- Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration
- Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data
- Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes
- Visual Language Models as Operator Agents in the Space Domain
- MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
- Geo-LLaVA: A Large Multi-Modal Model for Solving Geometry Math Problems with Meta In-Context Learning
- Cross-Attention Watermarking of Large Language Models
- IVCR-200K: A Large-Scale Multi-turn Dialogue Benchmark for Interactive Video Corpus Retrieval
- Zero-shot detection of buildings in mobile LiDAR using Language Vision Model
- Unsupervised Text Embedding Space Generation Using Generative Adversarial Networks for Text Synthesis
- DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
- Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
- Visual hallucination detection in large vision-language models via evidential conflict
- A Multimodal Symphony: Integrating Taste and Sound through Generative AI
- When Better Eyes Lead to Blindness: A Diagnostic Study of the Information Bottleneck in CNN-LSTM Image Captioning Models
- VILT: Video Instructions Linking for Complex Tasks
- Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
- Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
- On the Potential of CLIP for Compositional Logical Reasoning
- Level Up Your Tutorials: VLMs for Game Tutorials Quality Assessment
- Temporal Image Caption Retrieval Competition -- Description and Results
- Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval