Benchmark Radar
AI BENCHMARK PROFILE

KernelGenBench

General AIKnowledge & ReasoningFlagOS AI

KernelGenBench evaluates LLM- and agent-generated Triton kernels across 210 operators from three sources (ATen, vLLM, cuBLAS) and six hardware platforms, with automatic accuracy verification and two evaluation tracks (LLM Pass@K and iterative agent generation).

Released
2026-07-22
Readiness
Runnable
Primary field
General AI

Why it matters

Kernel generation is a specialized task lacking standardized evaluation; this benchmark provides a multi-source, multi-chip protocol to compare methods across diverse operators and hardware, enabling cost and portability assessment for autonomous kernel development.

Motivation

Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specialized and labor-intensive task.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.