Microsoft ResearchBITCambridgeMar 8, 2026arXiv:2603.07777

Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models

Zongqian Li, Shaohan Huang, Zewen Chi, Yixuan Su, Lexin Zhou, Li Dong, Nigel Collier, Furu Wei

AI Summary

The paper introduces MicroCoder-GRPO, a reinforcement learning method tailored for training code generation models, which incorporates conditional truncation masking, diversity-determined temperature selection, and KL loss removal with high clipping ratios. This approach addresses training bottlenecks arising from longer outputs and changing training dynamics in modern code generation models. Empirical results on LiveCodeBench v6 demonstrate up to 17.6% relative improvement over strong baselines, particularly under extended context evaluation, and the method is complemented by a new dataset and evaluator.

Key Contribution

By rethinking RLHF, MicroCoder-GRPO enables smaller code generation models to rival larger counterparts, achieving significant performance gains and revealing 34 training insights.

Abstract

Modern code generation models exhibit longer outputs, accelerated capability growth, and changed training dynamics, rendering traditional training methodologies, algorithms, and datasets ineffective for improving their performance. To address these training bottlenecks, we propose MicroCoder-GRPO, an improved Group Relative Policy Optimization approach with three innovations: conditional truncation masking to improve long output potential while maintaining training stability, diversity-determined temperature selection to maintain and encourage output diversity, and removal of KL loss with high clipping ratios to facilitate solution diversity. MicroCoder-GRPO achieves up to 17.6% relative improvement over strong baselines on LiveCodeBench v6, with more pronounced gains under extended context evaluation. Additionally, we release MicroCoder-Dataset, a more challenging training corpus that achieves 3x larger performance gains than mainstream datasets on LiveCodeBench v6 within 300 training steps, and MicroCoder-Evaluator, a robust framework with approximately 25% improved evaluation accuracy and around 40% faster execution. Through comprehensive analysis across more than thirty controlled experiments, we reveal 34 training insights across seven main aspects, demonstrating that properly trained models can achieve competitive performance with larger counterparts.

Code Generation & Program Synthesis RLHF & Preference Learning Training Efficiency & Optimization

Citation Metrics

Citations0

Influential citations0

References18

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models

Related Papers