AI BENCHMARK PROFILE
CodeGolf Bench
CodeGolf Bench evaluates concise code generation across 60 programming languages, using code golf platform problems and human performance baselines for scoring.
- Released
- 2026-05-28
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing code benchmarks focus on correctness, not efficiency or conciseness. CodeGolf Bench offers a unique measure of LLM ability to produce minimal solutions, complementing standard code generation evaluation.
Motivation
This paper introduces Code Bench, a benchmark capable of evaluating Large Language Models (LLMs) concise code generation abilities in 60 programming languages.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.