Benchmark Radar
AI BENCHMARK PROFILE

CodeGolf Bench

General AICoding & Software Engineering

CodeGolf Bench evaluates concise code generation across 60 programming languages, using code golf platform problems and human performance baselines for scoring.

Released
2026-05-28
Readiness
Paper only
Primary field
General AI

Why it matters

Existing code benchmarks focus on correctness, not efficiency or conciseness. CodeGolf Bench offers a unique measure of LLM ability to produce minimal solutions, complementing standard code generation evaluation.

Motivation

This paper introduces Code Bench, a benchmark capable of evaluating Large Language Models (LLMs) concise code generation abilities in 60 programming languages.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.