AI BENCHMARK PROFILE
ElephantBench
Closed-book knowledge probe with 1,094 questions evaluating recall of multiple divergent accounts for long-tail facts, with fixed C/P/F/K scoring metrics.
- Released
- 2026-08-28
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Provides a reproducible probe for diagnosing epistemic myopia in language models, with code and data available for direct use.
Motivation
Factual question answering (QA) typically assumes a single canonical answer, obscuring whether large language models (LLMs) retain divergent accounts of long-tail facts.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.