computer vision

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

arXiv:2607.26411

summary

The paper investigates whether unified multimodal models share a common semantic space by introducing cross-branch semantic steering, showing that semantic directions from the understanding branch can improve controllable image generation, while the reverse transfer is limited.

Abstract

Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabilities share a unified and transferable semantic space. This question is fundamentally challenging, as the two branches operate over heterogeneous representations (text tokens vs.\ visual latents) and distinct training objectives, making direct comparison difficult. To address this, we introduce \emph{cross-branch semantic steering}, an intervention-based framework that extracts semantic directions from one branch and applies them to the other. We show that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness. In contrast, the reverse direction consistently shows limited effectiveness. Our analysis suggests that this asymmetry may be related to a practical representational mismatch: understanding-derived vectors capture transferable, object-centric semantics, while generation-derived vectors primarily encode low-level appearance features. Our results reveal that architectural unification does not guarantee semantic alignment, and establish cross-branch steering as a practical tool for probing multimodal representations.

Topics & keywords

#unified multimodal models#semantic steering#cross-modal transfer#image synthesis#representation analysiscross-branch steeringsemantic directionsunderstanding branchgeneration branchobject-centric semantics
Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering · wovepaper