Search papers, labs, and topics across Lattice.
The paper addresses the structural blindness of Multimodal Large Language Models (MLLMs) when applied to engineering schematics by introducing a Vector-to-Graph (V2G) pipeline. V2G converts CAD diagrams into property graphs that explicitly represent component connectivity and dependencies. Experiments on an electrical compliance check benchmark demonstrate that V2G significantly improves accuracy compared to MLLMs, which struggle due to their pixel-driven approach.
MLLMs are structurally blind, but converting CAD diagrams to property graphs unlocks reliable schematic auditing.
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: structural blindness. Even state-of-the-art models fail to capture topology and symbolic logic in engineering schematics, as their pixel-driven paradigm discards the explicit vector-defined relations needed for reasoning. To overcome this, we propose a Vector-to-Graph (V2G) pipeline that converts CAD diagrams into property graphs where nodes represent components and edges encode connectivity, making structural dependencies explicit and machine-auditable. On a diagnostic benchmark of electrical compliance checks, V2G yields large accuracy gains across all error categories, while leading MLLMs remain near chance level. These results highlight the systemic inadequacy of pixel-based methods and demonstrate that structure-aware representations provide a reliable path toward practical deployment of multimodal AI in engineering domains. To facilitate further research, we release our benchmark and implementation at https://github.com/gm-embodied/V2G-Audit.