AI BENCHMARK PROFILE
PhantomFill
PhantomFill measures schema-coerced fabrication in language models by asking questions on unanswerable inputs under three output formats: free text, JSON with an escape option, and JSON with required fields. It reports Coerced Fabrication Rate and Escape Utilization Rate.
- Released
- 2026-06-11
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Hallucination in form-filling contexts is under-measured and costly. PhantomFill provides deterministic, code-based metrics targeting a critical failure mode, enabling model comparison and safety evaluation in structured output settings.
Motivation
Language models in production do not write prose.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.