HKUSTApr 14, 2026arXiv:2604.12268

CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation

Zaoyu Chen, Jianbo Dai, Boyu Zhu, Jingdong Wang, Huimin Wang, Huiming Wang, Haoyang Yuan, Zhijiang Guo, Xiao-Ming Wu

AI Summary

CodeSpecBench is introduced as a benchmark to evaluate LLMs' ability to generate executable behavioral specifications (pre/post conditions) for code, addressing limitations in existing evaluation methodologies. The benchmark includes function-level and repository-level tasks with specifications encoded as executable Python functions derived from real-world codebases. Evaluation of 15 LLMs reveals a significant performance drop on repository-level tasks, with the best model achieving only a 20.2% pass rate, and highlights that strong coding performance does not guarantee understanding of program semantics.

Key Contribution

LLMs that ace code generation often fail to grasp intended program semantics, as evidenced by a stark performance decline when generating executable behavioral specifications on the new CodeSpecBench benchmark.

Abstract

Large language models (LLMs) can generate code from natural language, but the extent to which they capture intended program behavior remains unclear. Executable behavioral specifications, defined via preconditions and postconditions, provide a concrete means to assess such understanding. However, existing work on specification generation is constrained in evaluation methodology, task settings, and specification expressiveness. We introduce CodeSpecBench, a benchmark for executable behavioral specification generation under an execution-based evaluation protocol. CodeSpecBench supports both function-level and repository-level tasks and encodes specifications as executable Python functions. Constructed from diverse real-world codebases, it enables a realistic assessment of both correctness (accepting valid behaviors) and completeness (rejecting invalid behaviors). Evaluating 15 state-of-the-art LLMs on CodeSpecBench, we observe a sharp performance degradation on repository-level tasks, where the best model attains only a 20.2% pass rate. We further find that specification generation is substantially more challenging than code generation, indicating that strong coding performance does not necessarily reflect deep understanding of intended program semantics. Our data and code are available at https://github.com/SparksofAGI/CodeSpecBench.

Code Generation & Program Synthesis Eval Frameworks & Benchmarks

Citation Metrics

Citations0

Influential citations0

References36

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation

Related Papers