NL2Scratch
NL2Scratch is an executable benchmark for natural-language-to-Scratch generation, consisting of 311,648 parser-valid NL-program pairs extracted from real Scratch projects. It includes a semantically validated pool of 23,594 examples and an 800-example diagnostic set. Evaluation uses Semantic Alignment Consistency (SAC), an interpretable slot-level metric for measuring semantic agreement.
- Released
- 2026-06-20
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing NL2Code evaluation focuses on text-based languages, leaving block-based programming unevaluated. NL2Scratch enables assessment of models on event-driven, visually compositional programs, revealing gaps between lexical similarity and semantic alignment that are invisible under token-level metrics.
Motivation
Block-based programming environments such as Scratch are widely used in early programming education, yet natural-language-to-code (NL2Code) research has focused primarily on text-based languages.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.