Benchmark Radar
AI BENCHMARK PROFILE

ElephantBench

General AIKnowledge & Reasoning

Closed-book knowledge probe with 1,094 questions evaluating recall of multiple divergent accounts for long-tail facts, with fixed C/P/F/K scoring metrics.

Released
2026-08-28
Readiness
Runnable
Primary field
General AI

Why it matters

Provides a reproducible probe for diagnosing epistemic myopia in language models, with code and data available for direct use.

Motivation

Factual question answering (QA) typically assumes a single canonical answer, obscuring whether large language models (LLMs) retain divergent accounts of long-tail facts.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.