Benchmark Radar
AI BENCHMARK PROFILE

Power Systems Agent Benchmark

General AIKnowledge & ReasoningPower Systems Agent Benchmark Team

Power Systems Agent Benchmark is an executable benchmark for power-engineering agents. It includes 41 task families with deterministic evaluators that recompute engineering quantities and check constraints, plus held-out generation.

Released
2026-06-18
Readiness
Runnable
Primary field
General AI

Why it matters

Executable evaluation in power engineering is missing. This benchmark provides a repeatable, contamination-resistant protocol for tool-using agents, with public and hidden splits.

Motivation

Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way to assess tool-using AI agents in software settings.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.