Search papers, labs, and topics across Lattice.
To predict downstream training dynamics across diverse model families, the authors construct "L-State," a framework combining static benchmark metrics with response vectors elicited by four short, target-independent training micro-interventions. This approach bridges a critical gap in model evaluation, proving that latent learning responsiveness can be reliably transferred across architectures via direct and structure-preserving operator readouts with explicit theoretical bounds. Across unseen evaluation targets, L-State readouts reduce response-prediction MSE by up to 78.3% on GLM-4-9B and halve RMSE on Granite-3.1-8B compared to evaluations relying on current capabilities alone.
A checkpoint's static benchmark score cannot predict how it will respond to further training, but four tiny probe interventions can forecast downstream fine-tuning trajectories across entirely unseen model families with up to 78% lower error.
Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. We measure this missing state by branching four short, standardized, target-independent micro-interventions from the same checkpoint and recording their effects in a common capability space. Together with current capability, these responses form L-State; its pulse block supports a flexible direct readout and a structure-preserving operator readout. Under smooth local dynamics, the operator construction admits an end-to-end cross-family bound with explicit source- and target-family coordinate heterogeneity. In three-family leave-one-family-out development, both pulse readouts reduce source-standardized MSE by 39.4% relative to capability alone, while separating the best response and direction estimates. On sealed GLM-4-9B, the direct and operator readouts reduce MSE by 71.8% and 78.3%, respectively, and the operator readout raises sign balanced accuracy from 0.366 to 0.754. On sealed Granite-3.1-8B, the direct readout reaches RMSE 0.544 and a development-fitted action-wise selector reaches 0.554, compared with 1.172 for capability alone. A five-family audit finds that the operator coordinate varies by action and family, and that modeling these deviations improves retrospective held-trajectory prediction. Target-independent interventions therefore expose training-response information that current capability misses, with direct and structured readouts covering complementary transfer regimes.