Benchmark Radar
AI BENCHMARK PROFILE

CodeAssay

General AICoding & Software Engineering

CodeAssay is a benchmark of 185 Python tasks across ten software-engineering categories with audited ground truth, public and hidden tests, and code-property measures. It evaluates LLM code generation correctness and other code properties.

Released
2026-08-04
Readiness
Paper only
Primary field
General AI

Why it matters

Provides a reproducible basis for evaluating LLM-generated code with validated ground truth and multiple metrics, enabling evidence-based model selection in software development.

Motivation

Large Language Models are increasingly evaluated for code generation using test-based benchmarks.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.