AI BENCHMARK PROFILE
SkillSafe-Bench
SkillSafe-Bench evaluates skill-merged LLMs on static refusal, adaptive jailbreak robustness, and capability retention using a two-judge AND rule, across multiple open-weight bases and attack types.
- Released
- 2026-08-09
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
It exposes that static safety does not predict robustness to adaptive attacks, providing a more accurate safety evaluation for model merging and guiding safer merging practices.
Motivation
Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.