AI BENCHMARK PROFILE
CodeAssay
CodeAssay is a benchmark of 185 Python tasks across ten software-engineering categories with audited ground truth, public and hidden tests, and code-property measures. It evaluates LLM code generation correctness and other code properties.
- Released
- 2026-08-04
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Provides a reproducible basis for evaluating LLM-generated code with validated ground truth and multiple metrics, enabling evidence-based model selection in software development.
Motivation
Large Language Models are increasingly evaluated for code generation using test-based benchmarks.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.