paper

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs

arXiv:2604.14520

Abstract

Omni-modal Large Language Models (Omni-MLLMs) are designed to reason over diverse sensory streams within a unified model. However, we find that multimodal reasoning is not solely determined by sensory evidence, but also by how modality information is organized. Through systematic analysis across diverse reasoning scenarios, we show that different modality organizations exhibit distinct advantages and limitations, and no single topology is universally optimal. Rather than proposing a general-purpose improvement to Omni-MLLM accuracy, we ask a narrower question: can these topology-induced failures be systematically identified and corrected? Motivated by this, we propose Chain of Modality (CoM), a framework that dynamically reorganizes multimodal topology during inference and learns adaptive organization strategies. Experiments across five benchmarks, diverse architectures, and model scales show that CoM reliably recovers topology-sensitive failures, through both training-free planning and lightweight Planner-SFT, while preserving performance on the full benchmark.