computer vision

FM: Unified Federated Foundation Models for Heterogeneous Multimodal Medical Imaging

arXiv:2607.13386

summary

The paper introduces FM², a federated learning framework that trains a unified foundation model for heterogeneous multimodal medical images while preserving privacy, using dual Mixture-of-Experts modules and caption-enhanced learning to align representations across institutions.

Abstract

Building foundation models for medical imaging requires pooling data across institutions, yet privacy regulations prohibit centralized aggregation. Existing Federated Foundation Models either fine-tune natural-image models with poor medical-domain transfer, or train from scratch within a single modality, lacking the flexibility to unify tasks. We identify an under-explored challenge, Imaging Modality Heterogeneity, where clients operate under two structural regimes: Overlapped (shared modalities with heterogeneous label distributions) and Non-overlapped (fully disjoint modalities per client). We propose FM, a unified framework that trains the core backbone from scratch to preserve medical domain fidelity while optionally incorporating biomedical pretrained encoders for vision-language alignment. FM equips each client with dual Mixture-of-Experts modules (a Class-wise MoE for personalized category knowledge and a Domain-wise MoE for shared cross-modality representations), coupled with a Heterogeneous Modality Alignment (HMA) regularizer that explicitly aligns modality-specific expert parameters, admitting provable convergence and generalization guarantees. FM further incorporates Caption-Enhanced Learning (CEL), where locally retained GPT-4o-generated captions serve as a textual semantic bridge enabling representation transfer across clients with disjoint modalities, and demonstrates extensibility to Federated Medical VQA. Experiments on our MIMH benchmark (classification and CEL) and real-world medical VQA datasets confirm consistent superiority over state-of-the-art federated baselines and strong out-of-modality generalization across all three tasks.

Accepted by ACM MM 2026 (Main Track): the 34th ACM International Conference on Multimedia

Topics & keywords

#federated learning#foundation models#multimodal medical imaging#privacy-preserving AI#heterogeneous modalitiesMixture-of-Expertsheterogeneous modality alignmentcaption-enhanced learningGPT-4o captionsO(1/√T) convergence
FM$^2$: Unified Federated Foundation Models for Heterogeneous Multimodal Medical Imaging · wovepaper