Search papers, labs, and topics across Lattice.
This study investigates how the format of theoretical content influences the programs generated by large language models (LLMs) in a randomized experiment involving 32 paired renderer slots. Despite the expectation that different presentation styles would yield distinct program outputs, the results showed no significant differences in program performance or identifiability across formats, with most evaluations failing to meet the predefined criteria. The findings challenge the assumption of uniform effects based on renderer format in LLM theory-to-program translation, establishing a boundary on the impact of specification formats in this context.
Renderer format fails to significantly influence the outputs of LLMs, undermining assumptions about its role in theory-to-program translation.
A verbal theory does not run: translating it into an executable model requires choices about variables, interventions, and interactions. We tested whether the presentation of otherwise identical theoretical content systematically changes the programs produced by large language models. In a prospective preregistered randomized experiment, 16 independent assignment bits allocated 32 paired renderer slots between a structured intervention contract and connected prose. Two pinned LLM snapshots translated five anonymous theoretical accounts, yielding 320 preauthorized single-shot programs in a frozen sparse quadratic language. A deterministic evaluator measured atomic finite-difference responses (H1) and mixed interaction responses (H2). Two primary endpoints assessed cross-model matched-distance reduction and closed-set same-account identifiability, with exact randomization inference and 27 preregistered support criteria evaluated across signed-linear and magnitude-rank pipelines. Both H1 and H2 returned the registered verdict NOT_SUPPORTED; only 19 of 108 criterion evaluations passed. Same-account identifiability remained near chance (AUC 0.469-0.523, against a registered 0.80 threshold). One H2 matched-distance endpoint moved and survived multiplicity correction in the signed-linear pipeline, but the corresponding magnitude-rank result missed the registered effect-size floor, so it did not satisfy the joint support rule. Thus renderer format did not produce the uniform, family-invariant, classifiable behavioral geometry predicted in advance. The result places a concrete boundary on specification-format effects in LLM theory-to-program translation and provides a fully auditable randomized design, datasets, and software for studying executable formalization.