Search papers, labs, and topics across Lattice.
This study investigates the role of computational evaluators in AI-assisted item development, revealing how representation, structural reduction, and selection policies critically influence which items reach expert review. Through two in-silico studies involving 32,000 Big Five items, the authors demonstrate that despite broad semantic agreement, local differences in item wording can lead to significant variations in construct evidence and item survival. The findings highlight that the computational evaluator is not a neutral intermediary but an integral component of measurement design, affecting the quality and consistency of psychometric evaluations.
Identical item wording can yield vastly different psychometric outcomes, revealing hidden instabilities in AI-assisted item development.
Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are usually treated as technical preliminaries. Yet representation, structural reduction, and selection policy determine which items and evidence psychometricians ever receive. Across two linked in-silico studies of 32,000 selected Big Five items, we followed fixed source populations from semantic representation through structural evaluation and candidate-form construction. Broad agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence, different items survived, and intended attributes could disappear even as community correspondence improved. These sensitivities also differed across generated source populations. At the final review boundary, both eligibility policies filled every content cell in every evaluable form, yet they presented different wording. Across embedding configurations, inclusive primary forms shared a median of only 6 of 40 items, reflecting the total downstream consequence of changing representation across structural evidence and ranking. The apparent stability of global summaries and complete forms therefore concealed instability in the content reaching psychometricians. The computational evaluator is not neutral infrastructure between generation and expertise; it is an inspectable and revisable part of measurement design.