Benchmark Radar
AI BENCHMARK PROFILE

FuzzingBrain-Bench

CybersecurityKnowledge & Reasoning

Evaluates LLMs on discovering distinct crashes in 77 open-source software challenges using sanitizer-instrumented Docker harnesses and deterministic scoring.

Released
2026-08-25
Readiness
Runnable
Primary field
Cybersecurity

Why it matters

Shifts vulnerability evaluation from predefined targets to open-ended crash discovery, providing a realistic and reproducible measure of LLM bug-finding capability.

Motivation

Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.