SRE-Bench
SRE-Bench evaluates AI agents on reverse engineering of binaries compiled from 19 private real-world-scale C programs (average 16.9K LoC) with 44 anti-analysis primitives, yielding 262 binary instances and 1572 deterministically graded tasks.
- Released
- 2026-08-11
- Readiness
- Paper only
- Primary field
- Cybersecurity
Why it matters
Existing benchmarks for agentic cybersecurity miss either contamination control or realistic scale. SRE-Bench addresses this gap by combining private, real-world-scale binaries with deterministic grading, enabling reliable measurement of agent performance in binary analysis and highlighting the gap between source-code and binary security capabilities.
Motivation
AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequential to cybersecurity, including malware, firmware, and proprietary applications, is available only as binaries.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.