Benchmark Radar
AI BENCHMARK PROFILE

SkillSafe-Bench

General AISafety & Trustworthiness

SkillSafe-Bench evaluates skill-merged LLMs on static refusal, adaptive jailbreak robustness, and capability retention using a two-judge AND rule, across multiple open-weight bases and attack types.

Released
2026-08-09
Readiness
Paper only
Primary field
General AI

Why it matters

It exposes that static safety does not predict robustness to adaptive attacks, providing a more accurate safety evaluation for model merging and guiding safer merging practices.

Motivation

Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.