Benchmark Radar
AI BENCHMARK PROFILE

CrackedPDFs

General AIKnowledge & Reasoning

Evaluates hidden prompt injection detection in PDFs through classification and paired ranking tasks, using 29,322 generated PDFs from 4,983 base documents. Includes frozen splits, features, and metrics for reproducibility.

Released
2026-07-03
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the gap in evaluating defenses against prompt injections embedded in PDF structure, where flattening can hide malicious instructions. Provides a controlled paired benchmark with confounding controls to assess whether detectors generalize beyond superficial cues.

Motivation

Document-based LLM systems often flatten a PDF before guardrails inspect it.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.