AI BENCHMARK PROFILE
SkillMisevo-Bench
Lifecycle-aware benchmark for persistent safety failures in self-improving LLM agents, consisting of 25 episodes with malicious demonstrations, benign twins, and fresh-session probes, scored with nine metrics including carryover ASR and unsafe retrieval.
- Released
- 2026-08-13
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Exposes how unsafe experiences can become reusable policies that cause later harm, enabling measurement of risk across skill authoring, retrieval, and reuse.
Motivation
Self-improving LLM agents convert successful trajectories into persistent cross-task state.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.