Benchmark Radar
AI BENCHMARK PROFILE

Boundary-Bench

General AIKnowledge & ReasoningBoundary-Bench authors

Boundary-Bench is an open-source plugin that adds configurable security policy levels to Terminal-Bench, enabling evaluation of coding agents under constraints like scoped credentials, restricted egress, read-only filesystems, and non-root execution.

Released
2026-08-02
Readiness
Paper only
Primary field
General AI

Why it matters

Existing coding agent benchmarks assume permissive sandboxes, leaving a gap in understanding performance under real-world security policies. This benchmark provides a standardized way to measure success and efficiency trade-offs across policy levels, informing model selection for deployment in hardened environments.

Motivation

Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems, non-root execution) constrain them like any other software.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.