Boundary-Bench
Boundary-Bench is an open-source plugin that adds configurable security policy levels to Terminal-Bench, enabling evaluation of coding agents under constraints like scoped credentials, restricted egress, read-only filesystems, and non-root execution.
- Released
- 2026-08-02
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing coding agent benchmarks assume permissive sandboxes, leaving a gap in understanding performance under real-world security policies. This benchmark provides a standardized way to measure success and efficiency trade-offs across policy levels, informing model selection for deployment in hardened environments.
Motivation
Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems, non-root execution) constrain them like any other software.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.