Search papers, labs, and topics across Lattice.
This study investigates how the output format of instruction-tuning data affects both data quality assessments and model performance benchmarks across various tasks and model families. By analyzing gradient signatures and employing spectral statistics, the authors reveal that the interface not only obscures the true quality signal but also significantly alters the perceived capability of tuned models, with performance gains being highly dependent on the training format. The findings indicate that current evaluation practices may misrepresent model capabilities by focusing on interface characteristics rather than the underlying content, leading to potential misinterpretations of model performance.
Output format can distort perceived model capabilities, with a 40-point accuracy gain in one format vanishing in another.
Instruction-tuning data are judged by quality metrics, and tuned models are judged by benchmarks, but both judgments pass through an output interface: the surface format in which an answer is written. Using gradient signatures across 12 tasks, four semantically equivalent interfaces, three model families, and controlled corruptions, we show that this interface confounds both measurements. Spectral statistics such as effective rank are provably invariant to interface rotation and empirically blind to semantic corruption, while the direction of the update carries the quality signal. The interface-varying residual is not noise: it identifies each unit's own target task perfectly across all three families. Capability itself is stored relative to the training interface: a skill that raises accuracy by more than 40 points under the training format can be nearly invisible under every other, and correcting a single generation budget flips the measured effect of fine-tuning on GSM8K from a gain into a large loss. Pre-registered interventions delimit where this geometry stops short of control. Data quality and model capability are interface-conditioned quantities, and current practice often reports the interface instead of the content.