Search papers, labs, and topics across Lattice.
This study investigates whether unified multimodal models (UMMs) operate within a single, transferable semantic space by employing a novel framework called cross-branch semantic steering. By extracting semantic directions from the understanding branch and applying them to the generation branch, the authors demonstrate that steering vectors from understanding enhance controllable image synthesis and semantic fidelity, while the reverse application yields limited results. The findings indicate that architectural unification does not ensure semantic alignment, highlighting a representational mismatch between object-centric semantics and low-level appearance features in UMMs.
Steering vectors from understanding to generation in unified multimodal models can significantly enhance image synthesis, but the reverse direction fails to deliver similar benefits.
Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabilities share a unified and transferable semantic space. This question is fundamentally challenging, as the two branches operate over heterogeneous representations (text tokens vs.\ visual latents) and distinct training objectives, making direct comparison difficult. To address this, we introduce \emph{cross-branch semantic steering}, an intervention-based framework that extracts semantic directions from one branch and applies them to the other. We show that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness. In contrast, the reverse direction consistently shows limited effectiveness. Our analysis suggests that this asymmetry may be related to a practical representational mismatch: understanding-derived vectors capture transferable, object-centric semantics, while generation-derived vectors primarily encode low-level appearance features. Our results reveal that architectural unification does not guarantee semantic alignment, and establish cross-branch steering as a practical tool for probing multimodal representations.