Benchmark Radar
AI BENCHMARK PROFILE

CodegenBench

General AICoding & Software Engineering

CodegenBench evaluates LLM-generated parallel code across x86_64, Sunway, and Kunpeng architectures using BLAS routines and specialized kernels. Scoring measures efficiency on each platform.

Released
2026-06-01
Readiness
Inspectable
Primary field
General AI

Why it matters

Addresses the gap in evaluating code generation for CPU-oriented HPC platforms, revealing cross-platform generalization limitations that matter for deploying LLMs in supercomputing contexts.

Motivation

While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.g., PyTorch, CUDA), their capabilities in CPU-oriented high-performance computing (HPC) across diverse architectures remain underexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.