AI BENCHMARK PROFILE
HoF-Bench
Evaluates vulnerability discovery in source code using 95 real CVEs across 8 repositories pinned at vulnerable commits, with a detector-blinded judge crediting findings that match code path, root cause, attack condition, and impact.
- Released
- 2026-07-29
- Readiness
- Runnable
- Primary field
- Cybersecurity
Why it matters
Provides a realistic test bed for comparing vulnerability scanners on rediscovery of known real-world vulnerabilities, with a strict scoring protocol and reusable dataset.
Motivation
LLM-based analyzers have begun finding real vulnerabilities in mature open-source projects: AISLE's analyzer is credited with more than 280 CVEs across 78 projects, including OpenSSL, curl, and GnuTLS.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.