A Survey on Multimodal Large Language Models
arXiv:2306.13549 · doi:10.1093/nsr/nwae403
Abstract
Recently, Multimodal Large Language Model (MLLM) represented by GPT-4V has been a new rising research hotspot, which uses powerful Large Language Models (LLMs) as a brain to perform multimodal tasks. The surprising emergent capabilities of MLLM, such as writing stories based on images and OCR-free math reasoning, are rare in traditional multimodal methods, suggesting a potential path to artificial general intelligence. To this end, both academia and industry have endeavored to develop MLLMs that can compete with or even better than GPT-4V, pushing the limit of research at a surprising speed. In this paper, we aim to trace and summarize the recent progress of MLLMs. First of all, we present the basic formulation of MLLM and delineate its related concepts, including architecture, training strategy and data, as well as evaluation. Then, we introduce research topics about how MLLMs can be extended to support more granularity, modalities, languages, and scenarios. We continue with multimodal hallucination and extended techniques, including Multimodal ICL (M-ICL), Multimodal CoT (M-CoT), and LLM-Aided Visual Reasoning (LAVR). To conclude the paper, we discuss existing challenges and point out promising research directions. In light of the fact that the era of MLLM has only just begun, we will keep updating this survey and hope it can inspire more research. An associated GitHub link collecting the latest papers is available at https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models.
Accepted for publication in National Science Review. Project page:https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models
References in corpus (48)
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- A Survey of Large Language Models
- Scaling Instruction-Finetuned Language Models
- LAION-5B: An open large-scale dataset for training next generation image-text models
- Reproducible scaling laws for contrastive language-image learning
- Fine-Tuning Language Models from Human Preferences
- nocaps: novel object captioning at scale
- Instruction Tuning with GPT-4
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
- LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
- Woodpecker: Hallucination Correction for Multimodal Large Language Models
- VideoChat: Chat-Centric Video Understanding
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
- Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models
- MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
- TALM: Tool Augmented Language Models
- MIMIC-IT: Multi-Modal In-Context Instruction Tuning
- Evaluation and Analysis of Hallucination in Large Vision-Language Models
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
- ImageBind-LLM: Multi-modality Instruction Tuning
- On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving
- Caption Anything: Interactive Image Description with Diverse Multimodal Controls
- mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
- Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
- Efficient Multimodal Learning from Data-centric Perspective
- TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
- AppAgent: Multimodal Agents as Smartphone Users
- ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst
- Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
- Chain of Thought Prompt Tuning in Vision Language Models
- MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
- Visual Programming: Compositional visual reasoning without training
- An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
- Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
- ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
- What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
- To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
- Silkie: Preference Distillation for Large Visual Language Models
Cited by in corpus (42)
- Unleashing the potential of prompt engineering for large language models
- Woodpecker: Hallucination Correction for Multimodal Large Language Models
- A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- Addressing Bias in Generative AI: Challenges and Research Opportunities in Information Management
- Vision-Language Models for Edge Networks: A Comprehensive Survey
- From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- Efficient Multimodal Large Language Models: A Survey
- Large language models for automated scholarly paper review: A survey
- Human-like object concept representations emerge naturally in multimodal large language models
- Reasoning Beyond Limits: Advances and Open Problems for LLMs
- A critical review of methods and challenges in large language models
- Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks
- Chat to Chip: Large Language Model Based Design of Arbitrarily Shaped Metasurfaces
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Leveraging Multimodal LLM for Inspirational User Interface Search
- Large Multimodal Models for Low-Resource Languages: A Survey
- LLM-Assisted Visual Analytics: Opportunities and Challenges
- Unleashing The Power of Pre-Trained Language Models for Irregularly Sampled Time Series
- Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions
- EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
- Audio-Language Datasets of Scenes and Events: A Survey
- Integrating Cognitive Processing Signals into Language Models: A Review of Advances, Applications and Future Directions
- Smart Glasses for CVI: Co-Designing Extended Reality Solutions to Support Environmental Perception by People with Cerebral Visual Impairment
- SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
- Can Large Language Models Grasp Concepts in Visual Content? A Case Study on YouTube Shorts about Depression
- LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- Reflections on "Can AI Understand Our Universe?"
- From Geometric Labels to Semantic Understanding of Indoor Building Components Using Multimodal Large Language Models
- FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting
- Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework
- Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
- Large Language Models -- the Future of Fundamental Physics?
- Seeing Red, Thinking Bad: Color Bias in Vision Language Models
- Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs
- Advanced Assistance for Traffic Crash Analysis: An AI-Driven Multi-Agent Approach to Pre-Crash Reconstruction
- Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs
- StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset
- When Visual Privacy Protection Meets Multimodal Large Language Models