Search papers, labs, and topics across Lattice.
This study investigates the use of synthetic semantic supervision for contrastive pretraining of small transformer encoders in code representation learning, addressing the limitations of human-written docstrings and costly execution traces. By employing a dual-encoder framework with synthetically generated natural-language descriptions that emphasize code functionality and intent, the authors benchmark their approach against various baselines across multiple programming languages. The results show significant performance improvements in five out of eight tasks compared to traditional pretraining methods, demonstrating that this approach can effectively compete with larger models while maintaining efficiency in training and inference.
Synthetic semantic supervision allows small transformers to outperform larger models in code representation tasks, offering a scalable alternative to traditional methods.
General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (setting-specific and costly to collect). We empirically study an alternative: contrastive pretraining of small encoders with synthetically generated natural-language descriptions emphasizing code functionality and intent, paired with code in a dual-encoder framework at training and discarded at inference. We benchmark this approach against pretraining-based baselines, generalist LLMs, and embedding-specific models on eight retrieval, classification, and generation tasks across C, C++, and Java. Synthetic semantic supervision yields statistically significant gains over pretraining baselines of the same inference-time size on five of eight tasks, with parity on two more; once fine-tuned, it matches or exceeds zero-shot models two orders of magnitude larger on classification, and it stays on par with execution-aware supervision at matched pretraining data, suggesting a scalable, effective alternative to existing code-representation paradigms.