Benchmark Radar
AI BENCHMARK PROFILE

SolidityBench

General AICoding & Software EngineeringSolidityBench Team

SolidityBench contains 5,470 repository-level Solidity smart contracts with natural language descriptions, plus SolidityScore, a semantic metric emphasizing domain-critical constructs. It evaluates code generation models.

Released
2026-06-18
Readiness
Paper only
Primary field
General AI

Why it matters

Domain-specific code generation lacks benchmarks. SolidityBench provides a resource to measure structural and semantic correctness in high-stakes smart contracts.

Motivation

Large Language Models (LLMs) have shown strong capabilities in general-purpose code generation, but their effectiveness in specialized software domains remains underexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.