Benchmark Radar
AI BENCHMARK PROFILE

AdvPlan-Bench

General AIKnowledge & Reasoning

AdvPlan-Bench is an offline benchmark for adversarial evaluation of structured plan-generation agents. It uses typed action chains, adversarial response sets, selector diagnostics, and metrics like BLUE-vs-RED advantage and Nash-gap. Includes 150 synthetic scenarios across five planning templates.

Released
2026-08-01
Readiness
Paper only
Primary field
General AI

Why it matters

Plan quality is often evaluated in isolation, but realistic tasks require considering adversarial responses. AdvPlan-Bench provides a reproducible method to study adversarial plan evaluation, response-budget sensitivity, and candidate frontiers, informing robust planning agent design.

Motivation

Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a candidate behaves when another agent can search for responses.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.