AI BENCHMARK PROFILE
Power Systems Agent Benchmark
Power Systems Agent Benchmark is an executable benchmark for power-engineering agents. It includes 41 task families with deterministic evaluators that recompute engineering quantities and check constraints, plus held-out generation.
- Released
- 2026-06-18
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Executable evaluation in power engineering is missing. This benchmark provides a repeatable, contamination-resistant protocol for tool-using agents, with public and hidden splits.
Motivation
Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way to assess tool-using AI agents in software settings.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.