AI BENCHMARK PROFILE
XPlainVerse
XPlainVerse evaluates deepfake detection and explanation quality, pairing real images with forgeries from twelve models and providing technical and simplified explanations, with metrics EntityScore and EvidenceScore for reasoning fidelity.
- Released
- 2026-07-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks focus on classification accuracy, not explanation grounding. XPlainVerse aims to measure whether explanations are grounded in actual manipulations, which is key for trustworthy deployable detection.
Motivation
As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limiting reliability and user trust.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.