Benchmark Radar
AI BENCHMARK PROFILE

PhantomFill

General AIKnowledge & Reasoning

PhantomFill measures schema-coerced fabrication in language models by asking questions on unanswerable inputs under three output formats: free text, JSON with an escape option, and JSON with required fields. It reports Coerced Fabrication Rate and Escape Utilization Rate.

Released
2026-06-11
Readiness
Runnable
Primary field
General AI

Why it matters

Hallucination in form-filling contexts is under-measured and costly. PhantomFill provides deterministic, code-based metrics targeting a critical failure mode, enabling model comparison and safety evaluation in structured output settings.

Motivation

Language models in production do not write prose.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.