KernelGenBench
KernelGenBench evaluates LLM- and agent-generated Triton kernels across 210 operators from three sources (ATen, vLLM, cuBLAS) and six hardware platforms, with automatic accuracy verification and two evaluation tracks (LLM Pass@K and iterative agent generation).
- Released
- 2026-07-22
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Kernel generation is a specialized task lacking standardized evaluation; this benchmark provides a multi-source, multi-chip protocol to compare methods across diverse operators and hardware, enabling cost and portability assessment for autonomous kernel development.
Motivation
Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specialized and labor-intensive task.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.