Benchmark Radar
AI BENCHMARK PROFILE

PERFOPT-Bench

General AICoding & Software Engineering

PERFOPT-Bench evaluates coding agents on software performance optimization tasks, requiring profiling, diagnosing bottlenecks, editing code while preserving correctness, and verifying reproducible speedups. Scoring includes hidden correctness tests, verified-speedup measurement, and trajectory-level audit.

Released
2026-07-08
Readiness
Inspectable
Primary field
General AI

Why it matters

Fills the gap in benchmarks focusing on performance engineering rather than functional correctness, measuring practical speedups on real execution targets and addressing issues like shortcut exploitation.

Motivation

Coding-agent benchmarks have largely measured whether agents can produce functionally correct patches, but production software also demands measurable speedups on real execution targets.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.