Benchmark Radar
AI BENCHMARK PROFILE

SkillMisevo-Bench

General AIKnowledge & ReasoningHenry Mao and collaborators

Lifecycle-aware benchmark for persistent safety failures in self-improving LLM agents, consisting of 25 episodes with malicious demonstrations, benign twins, and fresh-session probes, scored with nine metrics including carryover ASR and unsafe retrieval.

Released
2026-08-13
Readiness
Runnable
Primary field
General AI

Why it matters

Exposes how unsafe experiences can become reusable policies that cause later harm, enabling measurement of risk across skill authoring, retrieval, and reuse.

Motivation

Self-improving LLM agents convert successful trajectories into persistent cross-task state.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.