Search papers, labs, and topics across Lattice.
This paper introduces the Dual-prior Activation Residual Task-vectors Injection mechanism (DART-I), which enhances the reasoning capabilities of Multimodal Large Language Models (MLLMs) in interior design by addressing modality misalignment. By extracting spatial and color typography features and transforming them into directional task vectors, DART-I injects these vectors into the latent space of frozen MLLMs, allowing for precise reasoning without the need for fine-tuning. Experimental results show that DART-I significantly reduces hallucinations and improves spatial coherence in design outputs, demonstrating its effectiveness in constrained environments.
MLLMs can achieve precise interior design reasoning without fine-tuning, drastically reducing hallucinations and spatial errors.
Multimodal Large Language Models (MLLMs) have demonstrated great performance, yet they often suffer from severe modality misalignment when confronted with densely constrained spaces for interior design. Due to the loss of high-frequency local topological details and fine-grained aesthetic shifts during visual encoding, existing MLLMs frequently hallucinate, yielding physical spatial collisions and visual aesthetic dissonance. To address this, we propose Dual-prior Activation Residual Task-vectors Injection mechanism (DART-I) for MLLMs. It shifts the paradigm from lossy text-prompting to direct latent intervention, utilizing weak-to-strong deterministic rules to anchor the causal reasoning of MLLMs for interior design. Specifically, DART-I operates in three steps: it first explicitly extracts continuous spatial distance and color typography features from images using extremely lightweight weak experts; subsequently, it transforms these deterministic priors into directional task vectors via a linear projection network; these vectors are dynamically injected as residual terms into the latent space of the frozen MLLMs, steering MLLMs towards precise reasoning for interior design. Stepping outside the conventional paradigms, our method achieves precise reasoning without fine-tuning the MLLMs, effectively bypassing expensive computational costs and catastrophic forgetting. Extensive experiments on various benchmarks demonstrate the effectiveness and advantages of DART-I.