AI BENCHMARK PROFILE
PhantomBench
PhantomBench evaluates language models' ability to abstain from answering about non-existent entities. It comprises over 60,000 non-existent terms derived from real concepts across domains, and provides a pipeline for generating further instances.
- Released
- 2026-06-09
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the evaluation gap in assessing models' calibration of knowledge boundaries, offering a practical tool for detecting hallucination tendencies in high-stakes applications and studying behavior on rare concepts.
Motivation
Hallucinations, where language models (LMs) generate factually ungrounded responses, pose serious risks, as users tend to blindly rely on them.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.