SABER
SABER evaluates operational safety of LLM coding agents in stateful project workspaces. Agents perform realistic tasks, and safety is scored from the final environment state after action sequences. Violations are categorized by cause, enabling model-specific safety profile analysis.
- Released
- 2026-05-31
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing safety benchmarks only check prompt refusal, missing the impact of action sequences on workspaces. SABER fills this gap by measuring environment-aware operational safety, offering a practical way to compare models on their ability to avoid harmful state changes in realistic coding tasks.
Motivation
Large language models are increasingly deployed as coding agents, shifting safety from individual responses to action sequences.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.