Search papers, labs, and topics across Lattice.
This study investigates the mechanisms of knowledge transfer in heterogeneous multimodal large language models (MLLMs) using a novel approach called Cross-Scale Directional Parameter Injection (CDPI). By analyzing the transfer of capabilities across different model scales, the authors find that knowledge transfer is selective, with significant improvements in high-level reasoning but minimal gains in perceptual tasks. The results suggest that effective transfer occurs primarily in a low-ratio regime, challenging the assumption of broad capability inheritance in model fusion.
Knowledge transfer in MLLM fusion is not a blanket inheritance but a selective process favoring high-level reasoning over perception.
Training-free fusion of heterogeneous multimodal large language models (MLLMs) provides a direct route for cross-scale capability transfer, yet improvements in aggregate performance do not reveal what a smaller model actually inherits. Existing studies are largely designed and evaluated on limited task sets or aggregate metrics; as evaluation expands to broader task collections, whether different capabilities can transfer across scales remains poorly understood. To investigate this question, we introduce Cross-Scale Directional Parameter Injection (CDPI), a simple linear probe to analyze cross-scale knowledge transfer during heterogeneous fusion. A local theoretical analysis indicates that knowledge transfer selectivity is determined at first order by capability-dependent responses to a shared injection direction, while second-order curvature effects constrain the effective transfer regime. Across four Qwen3-VL model pairs and twelve multimodal benchmarks, our experiments reveal a consistent pattern of selectivity: gains concentrate on reasoning, particularly high-level reasoning, whereas perception performance remains close to that of the original target model. Component-wise ablations further show that high-level reasoning gains arise primarily from the language model, while ratio analysis finds that positive selective transfer occurs mainly in the small-ratio regime. These findings recast cross-scale heterogeneous MLLM fusion as selective language-side reasoning transfer within a narrow, low-interference regime, rather than broad capability inheritance.