Search papers, labs, and topics across Lattice.
This study introduces SciConsolidate, a method that transforms verified runtime experiences from scientific-computing tasks into transferable procedural knowledge, addressing the challenges of encoding source-specific repairs and bridging the abstraction-execution gap. By contrasting successes and failures, the method effectively induces cross-task procedures and utilizes failure-informed query synthesis to enhance consolidation data. The results show significant improvements in model performance, particularly in a stronger model context, highlighting the effectiveness of procedural guidance in overcoming execution limitations in scientific computing tasks.
Runtime procedure injection can boost model performance by nearly 7 points, revealing a critical gap between abstraction and execution in scientific computing.
Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems. We study scientific-computing experience consolidation: converting verified runtime experience into transferable procedural knowledge and persistent model improvement. This setting presents two challenges: trajectory-derived artifacts may encode source-specific repairs rather than cross-task computational mechanisms; and a weaker target model may be unable to operationalize an otherwise valid abstract procedure - an abstraction-execution gap. We introduce SciConsolidate, which contrasts verified successes and failures to induce cross-task procedures, selects them through a development-validation gate, and uses failure-informed, answer-free query synthesis to expand the consolidation data without requiring pre-existing reference answers. Because the target model may not directly execute these abstractions, a stronger model concretizes them into executable code supervision for standard, procedure-free SFT; a matched no-procedure teacher branch isolates the value of procedural guidance. On SciCode, runtime procedure injection improves Qwen3.6-27B by +3.85/+6.26 sub-step/main-problem points, but yields almost no aggregate main-problem gain for Qwen3.5-9B, providing operational evidence of the abstraction-execution gap. After procedure-guided concretization, the 9B student improves under procedure-free deployment by +3.89/+6.25 points over the no-procedure SFT control and by +5.62/+11.25 over the original 9B model. These results establish an experience-to-capability pathway for scientific computing and provide a practical starting point for scaling self-improving scientific assistance.