Search papers, labs, and topics across Lattice.
This paper introduces OraclePhys, a systematic framework for fine-tuning large language models (LLMs) specifically for structural mechanics, utilizing a novel benchmark and dataset that eliminate human labeling. The key findings reveal that the form of the answer label significantly influences the model's learning outcomes, with different objectives leading to varying levels of performance and generalization across tasks. Notably, the trained 8B model surpasses existing LLMs in precision on structural response tasks, demonstrating the critical role of answer form in shaping model capabilities.
The form of answer labels, not just their quantity, fundamentally shapes what LLMs learn during fine-tuning, revealing a surprising causal relationship that could redefine training strategies.
What a language model internalizes from fine-tuning is usually diagnosed after the fact. We make it an experimental variable. OraclePhys is a systematic fine-tuning framework with three components: OraclePhys-Bench, an exactly-graded structural-mechanics benchmark whose finite-element oracle scores every answer and counterfactual edit -- no human labels, no LLM judging; OraclePhys-30K, a supervision dataset of seven answer forms over byte-identical structure descriptions; and a controlled training study across the seven forms and three verifier roles. The study yields two findings. First, the label's answer form -- not its bit count -- causally determines what fine-tuning teaches: a ranking objective installs an out-of-distribution forward model where the untrained base sits at the guessing prior, a scalar objective at best a partial one, a boolean nothing detectable; the vector-scalar gulf survives a second physics domain, a second model family, and a paraphrased evaluation surface. Second, written or score-filtered answers install this capability, while advantage-weighted scores (GRPO) raise reward yet leave the model statistically equivalent to its start on held-out physics -- within the recipes and budgets tested -- sufficing only for routing. The trained 8B -- the first LLM on spatial structural response -- reaches the task's data-precision frontier: above a frontier LLM at zero- and 32-shot, at a specialist's level. What the label spells out about the target computation is what fine-tuning teaches; what you train on is what you route.