Benchmark Radar
AI BENCHMARK PROFILE

SEC-bench

CybersecurityKnowledge & Reasoning

SEC-bench Pro measures long-horizon vulnerability discovery in software security. It includes 344 verified vulnerabilities across V8, SpiderMonkey, and Linux kernel, each paired with instructions for reproducing a working proof-of-concept. Grading uses an LLM-based judge to classify generated PoCs against vulnerable, fixed, and latest images.

Released
2026-05-26
Readiness
Runnable
Primary field
Cybersecurity

Why it matters

Finding real vulnerabilities requires reasoning across an entire codebase to produce a working PoC, a challenging task that is understudied. SEC-bench Pro provides a reproducible environment for evaluating long-horizon security capabilities, helping identify where models succeed and fail in realistic vulnerability hunting.

Motivation

Finding a real vulnerability in complicated systems is a challenging, long-horizon task that demands reasoning across an entire codebase to produce a working proof-of-concept (PoC).

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.