Benchmark Radar
AI BENCHMARK PROFILE

CyberMaskQA

CybersecurityKnowledge & Reasoning

The evaluation object is a dataset for privacy-aware cybersecurity question answering, covering key security domains with private entity labels. The main capability evaluated is QA accuracy and masking performance.

Released
2026-05-23
Readiness
Paper only
Primary field
Cybersecurity

Why it matters

The evaluation gap is the lack of context-rich datasets for privacy-preserving QA in cybersecurity, which hinders progress. This benchmark's value is in enabling controlled information disclosure and studying privacy-utility trade-offs for deployable models.

Motivation

Large language models (LLMs) are increasingly applied to cybersecurity question answering (QA) for critical tasks such as incident response and vulnerability analysis.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.