Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
arXiv:2504.02477 · doi:10.1016/j.inffus.2025.103652
Abstract
Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We adopt a task-oriented perspective to systematically review the applications and advancements of multimodal fusion methods and VLMs in the field of robot vision. For semantic scene understanding tasks, we categorize fusion approaches into encoder-decoder frameworks, attention-based architectures, and graph neural networks. Meanwhile, we also analyze the architectural characteristics and practical implementations of these fusion strategies in key tasks such as simultaneous localization and mapping (SLAM), 3D object detection, navigation, and manipulation. We compare the evolutionary paths and applicability of VLMs based on large language models (LLMs) with traditional multimodal fusion methods.Additionally, we conduct an in-depth analysis of commonly used datasets, evaluating their applicability and challenges in real-world robotic scenarios. Building on this analysis, we identify key challenges in current research, including cross-modal alignment, efficient fusion, real-time deployment, and domain adaptation. We propose future directions such as self-supervised learning for robust multimodal representations, structured spatial memory and environment modeling to enhance spatial intelligence, and the integration of adversarial robustness and human feedback mechanisms to enable ethically aligned system deployment. Through a comprehensive review, comparative analysis, and forward-looking discussion, we provide a valuable reference for advancing multimodal perception and interaction in robotic vision. A comprehensive list of studies in this survey is available at https://github.com/Xiaofeng-Han-Res/MF-RV.
27 pages, 11 figures. Accepted to Information Fusion. Final journal version: volume 126 (Part B), February 2026
References in corpus (63)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- PaLM: Scaling Language Modeling with Pathways
- Flamingo: a Visual Language Model for Few-Shot Learning
- Linformer: Self-Attention with Linear Complexity
- Gemini: A Family of Highly Capable Multimodal Models
- Visual Instruction Tuning
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- CenterFusion: Center-based Radar and Camera Fusion for 3D Object Detection
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- PaLM-E: An Embodied Multimodal Language Model
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
- Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- IPOD: Intensive Point-based Object Detector for Point Cloud
- Unifying Voxel-based Representation with Transformer for 3D Object Detection
- Local Feature Matching Using Deep Learning: A Survey
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- DataComp: In search of the next generation of multimodal datasets
- Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone
- VIMA: General Robot Manipulation with Multimodal Prompts
- DeepInteraction: 3D Object Detection via Modality Interaction
- RD-VIO: Robust Visual-Inertial Odometry for Mobile Augmented Reality in Dynamic Environments
- DialFRED: Dialogue-Enabled Agents for Embodied Instruction Following
- Multi-task Paired Masking with Alignment Modeling for Medical Vision-Language Pre-training
- OpenVLA: An Open-Source Vision-Language-Action Model
- 3D-LLM: Injecting the 3D World into Large Language Models
- PaLI-X: On Scaling up a Multilingual Vision and Language Model
- An Introduction to Vision-Language Modeling
- MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
- A social context-aware graph-based multimodal attentive learning framework for disaster content classification during emergencies: a benchmark dataset and method
- Enriched Music Representations with Multiple Cross-modal Contrastive Learning
- CLFT: Camera-LiDAR Fusion Transformer for Semantic Segmentation in Autonomous Driving
- Touch and Go: Learning from Human-Collected Vision and Touch
- VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation
- An Embodied Generalist Agent in 3D World
- 3D-VLA: A 3D Vision-Language-Action Generative World Model
- MoCoViT: Mobile Convolutional Vision Transformer
- EA-LSS: Edge-aware Lift-splat-shot Framework for 3D BEV Object Detection
- MuSe-GNN: Learning Unified Gene Representation From Multimodal Biological Graph Data
- BEVBert: Multimodal Map Pre-training for Language-guided Navigation
- PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
- Joint Self-Supervised and Supervised Contrastive Learning for Multimodal MRI Data: Towards Predicting Abnormal Neurodevelopment
- Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
- RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
- MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection
- Advances in Embodied Navigation Using Large Language Models: A Survey
- RCBEVDet++: Toward High-accuracy Radar-Camera Fusion 3D Perception Network
- Ovis: Structural Embedding Alignment for Multimodal Large Language Model
- InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
- Real-to-Sim Grasp: Rethinking the Gap between Simulation and Real World in Grasp Detection
- Unleashing HyDRa: Hybrid Fusion, Depth Consistency and Radar for Unified 3D Perception
- GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs
- LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
- SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection
- InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
- SpaRC: Sparse Radar-Camera Fusion for 3D Object Detection
- InfiniteWorld: A Unified Scalable Simulation Framework for General Visual-Language Robot Interaction
- Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
- Sparsh: Self-supervised touch representations for vision-based tactile sensing