Benchmark Radar
AI BENCHMARK PROFILE

PhantomBench

General AIKnowledge & Reasoning

PhantomBench evaluates language models' ability to abstain from answering about non-existent entities. It comprises over 60,000 non-existent terms derived from real concepts across domains, and provides a pipeline for generating further instances.

Released
2026-06-09
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the evaluation gap in assessing models' calibration of knowledge boundaries, offering a practical tool for detecting hallucination tendencies in high-stakes applications and studying behavior on rare concepts.

Motivation

Hallucinations, where language models (LMs) generate factually ungrounded responses, pose serious risks, as users tend to blindly rely on them.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.