Benchmark Radar
AI BENCHMARK PROFILE

AgentS4D

General AISafety & Trustworthiness

AgentS4D evaluates runtime safety of LLM-based workspace agents across a four-dimensional framework, with 328 risk-injected cases and seven lifecycle checkpoints, measuring unsafe behavior and evidence across six risk-entry sources and nine harms.

Released
2026-07-29
Readiness
Paper only
Primary field
General AI

Why it matters

Existing safety benchmarks focus on endpoints, missing risks that emerge during execution. This benchmark provides a structured way to assess agent safety across the lifecycle, showing that task completion does not imply safety and that testing one risk form can miss vulnerabilities.

Motivation

Large language model (LLM)-based workspace agents execute stateful, multi-step workflows across heterogeneous resources, external tools, and persistent state.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.