PERFOPT-Bench
PERFOPT-Bench evaluates coding agents on software performance optimization tasks, requiring profiling, diagnosing bottlenecks, editing code while preserving correctness, and verifying reproducible speedups. Scoring includes hidden correctness tests, verified-speedup measurement, and trajectory-level audit.
- Released
- 2026-07-08
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Fills the gap in benchmarks focusing on performance engineering rather than functional correctness, measuring practical speedups on real execution targets and addressing issues like shortcut exploitation.
Motivation
Coding-agent benchmarks have largely measured whether agents can produce functionally correct patches, but production software also demands measurable speedups on real execution targets.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.