Benchmark Radar
AI BENCHMARK PROFILE

EduPluginBench

General AICoding & Software Engineering

EduPluginBench is an executable benchmark for evaluating code-generation models on producing plugins that meet governed ecosystem requirements including least privilege, telemetry consent, provenance, and bounded failure. It uses 1,440 mutants and 120 clean references with staged admission levels P0-P4.

Released
2026-08-01
Readiness
Paper only
Primary field
General AI

Why it matters

Compilation and functional tests do not ensure compliance with security and governance constraints. EduPluginBench provides a staged admission method targeting these gaps, enabling assessment of generated plugins in governed environments, which is critical for safe deployment.

Motivation

Code-generation models can produce executable components, but compilation and functional tests do not establish compliance with least privilege, telemetry consent, provenance, privileged-write authority, lifecycle constraints, or bounded failure.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.