AI BENCHMARK PROFILE
CrackedPDFs
Evaluates hidden prompt injection detection in PDFs through classification and paired ranking tasks, using 29,322 generated PDFs from 4,983 base documents. Includes frozen splits, features, and metrics for reproducibility.
- Released
- 2026-07-03
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Addresses the gap in evaluating defenses against prompt injections embedded in PDF structure, where flattening can hide malicious instructions. Provides a controlled paired benchmark with confounding controls to assess whether detectors generalize beyond superficial cues.
Motivation
Document-based LLM systems often flatten a PDF before guardrails inspect it.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.