AI BENCHMARK PROFILE
KSAFE-MM
KSAFE-MM evaluates multimodal LLM safety in Korean contexts, with 12 models tested on general and culture-specific safety risks, including jailbreak-style textual queries paired with local visual cues.
- Released
- 2026-05-27
- Readiness
- Paper only
- Primary field
- Cybersecurity
Why it matters
Existing safety benchmarks are English-centric and ignore local cultural risks; KSAFE-MM provides a general-to-local pipeline for culturally grounded safety evaluation, revealing trade-offs between safety and over-refusal.
Motivation
Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.