Benchmark Radar
AI BENCHMARK PROFILE

SABER

General AIAgentsSSSR Lab

SABER evaluates operational safety of LLM coding agents in stateful project workspaces. Agents perform realistic tasks, and safety is scored from the final environment state after action sequences. Violations are categorized by cause, enabling model-specific safety profile analysis.

Released
2026-05-31
Readiness
Runnable
Primary field
General AI

Why it matters

Existing safety benchmarks only check prompt refusal, missing the impact of action sequences on workspaces. SABER fills this gap by measuring environment-aware operational safety, offering a practical way to compare models on their ability to avoid harmful state changes in realistic coding tasks.

Motivation

Large language models are increasingly deployed as coding agents, shifting safety from individual responses to action sequences.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.