PTXBench
PTXBench evaluates LLMs in generating architecture-specific PTX for GPU kernel optimization. It measures functional correctness, execution of target instructions, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs.
- Released
- 2026-08-18
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Addresses the lack of standardized evaluation for LLM-driven GPU kernel optimization with architecture-specific PTX, providing a reproducible testbed to compare model capabilities and guide improvements in exploiting evolving GPU architectures.
Motivation
We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.