AI BENCHMARK PROFILE
FuzzingBrain-Bench
Evaluates LLMs on discovering distinct crashes in 77 open-source software challenges using sanitizer-instrumented Docker harnesses and deterministic scoring.
- Released
- 2026-08-25
- Readiness
- Runnable
- Primary field
- Cybersecurity
Why it matters
Shifts vulnerability evaluation from predefined targets to open-ended crash discovery, providing a realistic and reproducible measure of LLM bug-finding capability.
Motivation
Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.