Benchmark Radar
AI BENCHMARK PROFILE

DeCompBench

General AISafety & Trustworthiness

DeCompBench evaluates agent safety against decomposition attacks, where harmful tasks are broken into benign subtasks; it measures refusal rates and objective fulfillment on decomposed variants.

Released
2026-06-12
Readiness
Inspectable
Primary field
General AI

Why it matters

Addresses a security gap not covered by existing agent safety benchmarks, critical for preventing adversarial misuse in deployed agents.

Motivation

LLM-based Agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real-world.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.