Search papers, labs, and topics across Lattice.
This paper introduces simple self-distillation (SSD), a method for improving LLM code generation by fine-tuning on the model's own samples generated with specific temperature and truncation. Applying SSD to Qwen3-30B-Instruct improves its pass@1 score on LiveCodeBench v6 from 42.4% to 55.3%, especially on harder problems. Analysis reveals that SSD resolves a precision-exploration conflict in LLM decoding by reshaping token distributions to suppress distractors and preserve diversity.
Self-distillation, using only an LLM's own raw outputs, can boost its code generation performance by over 10 points on pass@1, rivaling more complex methods.
Can a large language model (LLM) improve at code generation using only its own raw outputs, without a verifier, a teacher model, or reinforcement learning? We answer in the affirmative with simple self-distillation (SSD): sample solutions from the model with certain temperature and truncation configurations, then fine-tune on those samples with standard supervised fine-tuning. SSD improves Qwen3-30B-Instruct from 42.4% to 55.3% pass@1 on LiveCodeBench v6, with gains concentrating on harder problems, and it generalizes across Qwen and Llama models at 4B, 8B, and 30B scale, including both instruct and thinking variants. To understand why such a simple method can work, we trace these gains to a precision-exploration conflict in LLM decoding and show that SSD reshapes token distributions in a context-dependent way, suppressing distractor tails where precision matters while preserving useful diversity where exploration matters. Taken together, SSD offers a complementary post-training direction for improving LLM code generation.