AI BENCHMARK PROFILE
OBLIVION
Evaluates operational skill unlearning in deployed agents across 88 attack episodes, measuring attack success rate and impact-weighted exposure after workflow-level defenses.
- Released
- 2026-08-08
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Introduces a measurable benchmark for a new safety problem: preventing agents from rebuilding revoked skills, supporting workflow-level evaluation beyond parameter forgetting.
Motivation
Large language model agents are becoming operational interfaces to files, memories, registries, and external tools.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.