Search papers, labs, and topics across Lattice.
This paper introduces VEHBench, a diagnostic benchmark tailored for LLM-assisted design of vibration energy harvesters (VEHs), which evaluates LLM performance across four distinct design roles. By employing an analytical physical oracle to score 763 tasks, the study reveals that LLM capabilities vary significantly depending on the design stage, with no single model excelling throughout the entire process. These insights highlight the necessity for stage-aware evaluation frameworks in engineering workflows, enabling more effective selection and improvement of LLMs in complex design tasks.
LLM performance in vibration energy harvester design is stage-dependent, with no single model consistently outperforming others across the entire workflow.
Battery-free Internet of Things (IoT) requires iterative design of vibration energy harvesters (VEHs) under coupled physical constraints, while LLMs are emerging as interface layers for engineering workflows. However, existing engineering benchmarks primarily assess final artifact validity, offering limited insights into how LLMs behave across different stages of coupled physical design. We introduce VEHBench, an engineering-native diagnostic benchmark for LLM-assisted VEH design, featuring 763 literature-grounded tasks scored by an analytical physical oracle. VEHBench evaluates four design roles: specification triage, verifier-guided search, corrupted-state recovery, and policy-conditioned selection. Experimental results reveal that LLM capability is strongly stage-dependent: no single model consistently dominates the entire workflow, and response-control profiles expose distinct behavioral patterns across design roles. VEHBench thus provides a stage-aware foundation for evaluating, selecting, routing, and improving verifier-grounded engineering LLMs. The benchmark artifact is available at https://huggingface.co/datasets/AnonymousVehbench/vehbench