Benchmark Radar
AI BENCHMARK PROFILE

CyberChainBench

CybersecurityMultimodal Perception

CyberChainBench evaluates LLM-based agents on smart contract security across vulnerability detection, exploit generation, and patch synthesis, using 541 real exploit incidents with on-chain evaluation on historical forks.

Released
2026-06-24
Readiness
Paper only
Primary field
Cybersecurity

Why it matters

Smart contract security requires realistic, end-to-end evaluation of agent capabilities; CyberChainBench provides structured ground truth and economic impact metrics to measure practical effectiveness.

Motivation

We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability detection, exploit generation, and patch synthesis.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.