Benchmark Radar
AI BENCHMARK PROFILE

PTXBench

General AICoding & Software EngineeringPTXBench Team

PTXBench evaluates LLMs in generating architecture-specific PTX for GPU kernel optimization. It measures functional correctness, execution of target instructions, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs.

Released
2026-08-18
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the lack of standardized evaluation for LLM-driven GPU kernel optimization with architecture-specific PTX, providing a reproducible testbed to compare model capabilities and guide improvements in exploiting evolving GPU architectures.

Motivation

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.