CodegenBench
CodegenBench evaluates LLM-generated parallel code across x86_64, Sunway, and Kunpeng architectures using BLAS routines and specialized kernels. Scoring measures efficiency on each platform.
- Released
- 2026-06-01
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Addresses the gap in evaluating code generation for CPU-oriented HPC platforms, revealing cross-platform generalization limitations that matter for deploying LLMs in supercomputing contexts.
Motivation
While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.g., PyTorch, CUDA), their capabilities in CPU-oriented high-performance computing (HPC) across diverse architectures remain underexplored.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.