Search papers, labs, and topics across Lattice.
This paper introduces CommitLLM, a three-stage pipeline that generates concise and compliant git commit messages from code diffs using a fine-tuned small language model. By employing QLoRA fine-tuning on the Mistral-7B-Instruct-v0.2 model and implementing constrained decoding and deterministic post-processing, the system significantly enhances message quality and compliance. The results show a dramatic increase in format adherence from 22% to 98% and a reduction in average output length, highlighting the importance of structured output in improving LLM performance for specific tasks.
CommitLLM transforms uninformative git commit messages into concise, compliant formats, achieving 98% adherence and drastically reducing message length.
Developers frequently write uninformative git commit messages such as"fix"or"update stuff", degrading the value of version-control history for code review, debugging, and onboarding. We present CommitLLM, a three-stage pipeline that generates concise, Conventional Commits-compliant messages from code diffs using a fine-tuned small language model. The system combines (1) QLoRA fine-tuning of Mistral-7B-Instruct-v0.2 on the CommitPackFT dataset, (2) constrained decoding to enforce brevity, and (3) deterministic post-processing to strip conversational artifacts and enforce format. On a 50-sample evaluation, CommitLLM achieves 98% format compliance (vs. 22% for vanilla Mistral), reduces average output length from 154.8 to 37.9 characters, and improves LLM-as-a-Judge scores from 1.97 to 3.68 out of 5. Notably, the post-processing layers contribute more to quality improvement than the fine-tuning itself, suggesting that for structured-output tasks, treating the LLM as a component in a deterministic pipeline is more effective than optimizing the model alone. The entire system runs on a single consumer GPU (NVIDIA T4, 16 GB VRAM).