From Screens to Scenes: A Survey of Embodied AI in Healthcare
arXiv:2501.07468 · doi:10.1016/j.inffus.2025.103033
Abstract
Healthcare systems worldwide face persistent challenges in efficiency, accessibility, and personalization. Powered by modern AI technologies such as multimodal large language models and world models, Embodied AI (EmAI) represents a transformative frontier, offering enhanced autonomy and the ability to interact with the physical world to address these challenges. As an interdisciplinary and rapidly evolving research domain, "EmAI in healthcare" spans diverse fields such as algorithms, robotics, and biomedicine. This complexity underscores the importance of timely reviews and analyses to track advancements, address challenges, and foster cross-disciplinary collaboration. In this paper, we provide a comprehensive overview of the "brain" of EmAI for healthcare, wherein we introduce foundational AI algorithms for perception, actuation, planning, and memory, and focus on presenting the healthcare applications spanning clinical interventions, daily care & companionship, infrastructure support, and biomedical research. Despite its promise, the development of EmAI for healthcare is hindered by critical challenges such as safety concerns, gaps between simulation platforms and real-world applications, the absence of standardized benchmarks, and uneven progress across interdisciplinary domains. We discuss the technical barriers and explore ethical considerations, offering a forward-looking perspective on the future of EmAI in healthcare. A hierarchical framework of intelligent levels for EmAI systems is also introduced to guide further development. By providing systematic insights, this work aims to inspire innovation and practical applications, paving the way for a new era of intelligent, patient-centered healthcare.
56 pages, 11 figures, manuscript accepted by Information Fusion
References in corpus (155)
- Denoising Diffusion Probabilistic Models
- A Brief Survey of Deep Reinforcement Learning
- Training language models to follow instructions with human feedback
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- LoRA: Low-Rank Adaptation of Large Language Models
- On the Opportunities and Risks of Foundation Models
- Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
- A Survey of Large Language Models
- BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining
- Emergent Abilities of Large Language Models
- Gemini: A Family of Highly Capable Multimodal Models
- A Survey on Multimodal Large Language Models
- Visual Instruction Tuning
- ReAct: Synergizing Reasoning and Acting in Language Models
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- QLoRA: Efficient Finetuning of Quantized LLMs
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
- PaLM-E: An Embodied Multimodal Language Model
- Florence: A New Foundation Model for Computer Vision
- Multi-Modal Knowledge Graph Construction and Application: A Survey
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
- Reflexion: Language Agents with Verbal Reinforcement Learning
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
- BioRED: A Rich Biomedical Relation Extraction Dataset
- Rearrangement: A Challenge for Embodied AI
- Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping
- Visual Dexterity: In-Hand Reorientation of Novel and Complex Object Shapes
- DALL-E-Bot: Introducing Web-Scale Diffusion Models to Robotics
- AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
- ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models
- ProcTHOR: Large-Scale Embodied AI Using Procedural Generation
- AI Alignment: A Comprehensive Survey
- Spatial Action Maps for Mobile Manipulation
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- Enabling AI and Robotic Coaches for Physical Rehabilitation Therapy: Iterative Design and Evaluation with Therapists and Post-Stroke Survivors
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- Tele-operative Robotic Lung Ultrasound Scanning Platform for Triage of COVID-19 Patients
- Language Models as Zero-Shot Trajectory Generators
- Pre-Trained Language Models for Interactive Decision-Making
- Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
- Precise Repositioning of Robotic Ultrasound: Improving Registration-based Motion Compensation using Ultrasound Confidence Optimization
- OpenVLA: An Open-Source Vision-Language-Action Model
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought
- Large Language Models for Robotics: A Survey
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning
- MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
- Augmenting Language Models with Long-Term Memory
- Understanding Domain Randomization for Sim-to-real Transfer
- Is Conditional Generative Modeling all you need for Decision-Making?
- Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model
- Behavior Transformers: Cloning modes with one stone
- A Survey on Vision-Language-Action Models for Embodied AI
- The SARAS Endoscopic Surgeon Action Detection (ESAD) dataset: Challenges and methods
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
- Language Models Meet World Models: Embodied Experiences Enhance Language Models
- DISC-MedLLM: Bridging General Large Language Models and Real-World Medical Consultation
- Hit and Lead Discovery with Explorative RL and Fragment-based Molecule Generation
- Exploiting Natural Language for Efficient Risk-Aware Multi-robot SaR Planning
- SurgiTrack: Fine-Grained Multi-Class Multi-Tool Tracking in Surgical Videos
- Wider and Deeper LLM Networks are Fairer LLM Evaluators
- A Survey of Embodied Learning for Object-Centric Robotic Manipulation
- Modeling of Pruning Techniques for Deep Neural Networks Simplification
- An Embodied Generalist Agent in 3D World
- LLM Evaluators Recognize and Favor Their Own Generations
- FiLM-Ensemble: Probabilistic Deep Learning via Feature-wise Linear Modulation
- UniAudio: An Audio Foundation Model Toward Universal Audio Generation
- Large Language Models and Causal Inference in Collaboration: A Survey
- ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory
- Voice control interface for surgical robot assistants
- Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
- SmartArm: Suturing Feasibility of a Surgical Robotic System on a Neonatal Chest Model
- Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
- World Model on Million-Length Video And Language With Blockwise RingAttention
- Energy-Based Imitation Learning
- Sparks of Large Audio Models: A Survey and Outlook
- Cosmos World Foundation Model Platform for Physical AI
- VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
- A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
- Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
- AlphaBlock: Embodied Finetuning for Vision-Language Reasoning in Robot Manipulation
- A Review of Multi-Modal Large Language and Vision Models
- Surgical Text-to-Image Generation
- Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
- Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
- Think-in-Memory: Recalling and Post-thinking Enable LLMs with Long-Term Memory
- Faithful Explanations of Black-box NLP Models Using LLM-generated Counterfactuals
- Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery
- Where2Explore: Few-shot Affordance Learning for Unseen Novel Categories of Articulated Objects
- CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
- MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation
- Implicit Distributional Reinforcement Learning
- How Far is Video Generation from World Model: A Physical Law Perspective
- RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs
- RoboNurse-VLA: Robotic Scrub Nurse System based on Vision-Language-Action Model
- PlanAgent: A Multi-modal Large Language Agent for Closed-loop Vehicle Motion Planning
- Towards End-to-End Embodied Decision Making via Multi-modal Large Language Model: Explorations with GPT4-Vision and Beyond
- MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
- Exploring the In-context Learning Ability of Large Language Model for Biomedical Concept Linking
- NatSGD: A Dataset with Speech, Gestures, and Demonstrations for Robot Learning in Natural Human-Robot Interaction
- SurGen: Text-Guided Diffusion Model for Surgical Video Generation
- GP-GPT: Large Language Model for Gene-Phenotype Mapping
- EquiBot: SIM(3)-Equivariant Diffusion Policy for Generalizable and Data Efficient Learning
- Benchmarking and Explaining Large Language Model-based Code Generation: A Causality-Centric Approach
- 15M Multimodal Facial Image-Text Dataset
- Interactive Task Planning with Language Models
- TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
- Probabilistically Correct Language-based Multi-Robot Planning using Conformal Prediction
- Re-ReST: Reflection-Reinforced Self-Training for Language Agents
- Bora: Biomedical Generalist Video Generation Model
- GP-VLS: A general-purpose vision language model for surgery
- Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
- LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models
- SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
- MLLM-Search: A Zero-Shot Approach to Finding People using Multimodal Large Language Models
- Sensory Glove-Based Surgical Robot User Interface
- Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
- KARMA: Augmenting Embodied AI Agents with Long-and-short Term Memory Systems
- 3-Survivor: A Rough Terrain Negotiable Teleoperated Mobile Rescue Robot with Passive Control Mechanism
- StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
- Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models
- On the assesment of functional connectivity in an immersive brain-computer interface during motor imagery
- LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning
- Visual Motion Imagery Classification with Deep Neural Network based on Functional Connectivity
- SurgicAI: A Hierarchical Platform for Fine-Grained Surgical Policy Learning and Benchmarking
- Chemistry3D: Robotic Interaction Benchmark for Chemistry Experiments
- Medical Video Generation for Disease Progression Simulation
- Building an Affordances Map with Interactive Perception
- ProtCLIP: Function-Informed Protein Multi-Modal Learning
- A Medical Low-Back Pain Physical Rehabilitation Dataset for Human Body Movement Analysis
- Ultrasound-QBench: Can LLMs Aid in Quality Assessment of Ultrasound Imaging?
- Personalized Heart Disease Detection via ECG Digital Twin Generation
- Cybersecurity and Embodiment Integrity for Modern Robots: A Conceptual Framework
- Empowering Embodied Manipulation: A Bimanual-Mobile Robot Manipulation Dataset for Household Tasks
- SurgicalGS: Dynamic 3D Gaussian Splatting for Accurate Robotic-Assisted Surgical Scene Reconstruction
- Surgical Task Automation Using Actor-Critic Frameworks and Self-Supervised Imitation Learning
- Natural Language as Policies: Reasoning for Coordinate-Level Embodied Control with LLMs
- Learning to Generate Context-Sensitive Backchannel Smiles for Embodied AI Agents with Applications in Mental Health Dialogues
- ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models
- Text-centric Alignment for Multi-Modality Learning
- RoboGolf: Mastering Real-World Minigolf with a Reflective Multi-Modality Vision-Language Model
- MERCI: Multimodal Emotional and peRsonal Conversational Interactions Dataset
- Less is More: A Closer Look at Semantic-based Few-Shot Learning
- Conditional Prompt Tuning for Multimodal Fusion
- MAEA: Multimodal Attribution for Embodied AI
- Enhancing Deformable Object Manipulation By Using Interactive Perception and Assistive Tools