Visual Instruction Tuning
arXiv:2304.08485
Abstract
Instruction tuning large language models (LLMs) using machine-generated instruction-following data has improved zero-shot capabilities on new tasks, but the idea is less explored in the multimodal field. In this paper, we present the first attempt to use language-only GPT-4 to generate multimodal language-image instruction-following data. By instruction tuning on such generated data, we introduce LLaVA: Large Language and Vision Assistant, an end-to-end trained large multimodal model that connects a vision encoder and LLM for general-purpose visual and language understanding.Our early experiments show that LLaVA demonstrates impressive multimodel chat abilities, sometimes exhibiting the behaviors of multimodal GPT-4 on unseen images/instructions, and yields a 85.1% relative score compared with GPT-4 on a synthetic multimodal instruction-following dataset. When fine-tuned on Science QA, the synergy of LLaVA and GPT-4 achieves a new state-of-the-art accuracy of 92.53%. We make GPT-4 generated visual instruction tuning data, our model and code base publicly available.
NeurIPS 2023 Oral; project page: https://llava-vl.github.io/
Cited by in corpus (47)
- Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
- Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education
- Materials science in the era of large language models: a perspective
- A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
- From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
- CLIP in Medical Imaging: A Survey
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- CFN-ESA: A Cross-Modal Fusion Network with Emotion-Shift Awareness for Dialogue Emotion Recognition
- Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
- Emotion Detection for Misinformation: A Review
- "It's not like Jarvis, but it's pretty close!" -- Examining ChatGPT's Usage among Undergraduate Students in Computer Science
- When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
- A Survey on Occupancy Perception for Autonomous Driving: The Information Fusion Perspective
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- Exploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
- OpenECAD: An Efficient Visual Language Model for Editable 3D-CAD Design
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- LLaVA-Docent: Instruction Tuning with Multimodal Large Language Model to Support Art Appreciation Education
- Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
- Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks
- CogniAlign: Word-Level Multimodal Speech Alignment with Gated Cross-Attention for Alzheimer's Detection
- EmoVerse: Exploring Multimodal Large Language Models for Sentiment and Emotion Understanding
- RSTeller: Scaling Up Visual Language Modeling in Remote Sensing with Rich Linguistic Semantics from Openly Available Data and Large Language Models
- MMSR: Symbolic Regression is a Multi-Modal Information Fusion Task
- Tucano: Advancing Neural Text Generation for Portuguese
- IFShip: Interpretable Fine-grained Ship Classification with Domain Knowledge-Enhanced Vision-Language Models
- Towards an automated workflow in materials science for combining multi-modal simulative and experimental information using data mining and large language models
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
- Unified Modeling Language Code Generation from Diagram Images Using Multimodal Large Language Models
- Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
- Arabic Mini-ClimateGPT : A Climate Change and Sustainability Tailored Arabic LLM
- Visual Language Models as Operator Agents in the Space Domain
- Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes
- An Explainable Deep Neural Network with Frequency-Aware Channel and Spatial Refinement for Flood Prediction in Sustainable Cities
- iCub Detecting Gazed Objects: A Pipeline Estimating Human Attention
- FDM-Bench: A Comprehensive Benchmark for Evaluating Large Language Models in Additive Manufacturing Tasks
- ROSGPT_Vision: Commanding Robots Using Only Language Models' Prompts
- SituationalLLM: Proactive language models with scene awareness for dynamic, contextual task guidance
- Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
- Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success using Knowledge-infused Learning
- Visual hallucination detection in large vision-language models via evidential conflict
- Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization
- Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
- DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models