Search papers, labs, and topics across Lattice.
This paper introduces FuzzingBrain-Bench, a novel benchmark designed to evaluate the bug discovery capabilities of large language models (LLMs) on open-source software. Unlike traditional methods that focus on predefined vulnerabilities, this benchmark allows models to generate inputs that trigger a wide range of distinct crashes, providing a more comprehensive assessment of their capabilities. The evaluation reveals that Claude Opus 4.8 outperforms other models by successfully triggering crashes in 60 out of 77 challenges, highlighting the potential of LLMs in open-ended bug discovery.
LLMs can uncover a surprising variety of software bugs, with Claude Opus 4.8 achieving a notable success rate in a new benchmark designed for open-ended bug discovery.
Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the model to generate a proof-of-concept input that triggers a predefined target vulnerability. However, this setup may overlook valid crashes discovered by the model when they do not match the predefined target. As a result, the evaluation may not reflect the model's real capability. We present FuzzingBrain-Bench, a benchmark for assessing AI models'ability to discover bugs in open-source software. Models are given an open-source project and a sanitizer-instrumented harness in a self-contained Docker image. Their goal is to generate inputs that trigger as many distinct crashes as possible through the harness. A model's performance on each challenge is scored based on the number of distinct crash signatures it produces, capped at a predefined maximum and weighted by a difficulty coefficient. FuzzingBrain-Bench V1 consists of 77 challenges drawn from 43 open-source projects, with 36 C, 32 C++, and 9 Java/JVM challenges. We evaluate Claude Haiku 4.5, Claude Sonnet 4.6, and Claude Opus 4.8 on the full benchmark. Claude Opus 4.8 performs best, triggering crashes in 60 of 77 challenges and achieving a score of 196 out of 579. None of the three models triggers a crash in 13 challenges. The FuzzingBrain-Bench corpus and harnesses are publicly available at https://github.com/fuzzingbrain/FuzzingBrain-Bench.