Benchmark Radar
AI BENCHMARK PROFILE

OBLIVION

General AIKnowledge & ReasoningOBLIVION Team

Evaluates operational skill unlearning in deployed agents across 88 attack episodes, measuring attack success rate and impact-weighted exposure after workflow-level defenses.

Released
2026-08-08
Readiness
Paper only
Primary field
General AI

Why it matters

Introduces a measurable benchmark for a new safety problem: preventing agents from rebuilding revoked skills, supporting workflow-level evaluation beyond parameter forgetting.

Motivation

Large language model agents are becoming operational interfaces to files, memories, registries, and external tools.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.