AI BENCHMARK PROFILE
CyberChainBench
CyberChainBench evaluates LLM-based agents on smart contract security across vulnerability detection, exploit generation, and patch synthesis, using 541 real exploit incidents with on-chain evaluation on historical forks.
- Released
- 2026-06-24
- Readiness
- Paper only
- Primary field
- Cybersecurity
Why it matters
Smart contract security requires realistic, end-to-end evaluation of agent capabilities; CyberChainBench provides structured ground truth and economic impact metrics to measure practical effectiveness.
Motivation
We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability detection, exploit generation, and patch synthesis.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.