Benchmark Radar
AI BENCHMARK PROFILE

XPlainVerse

General AIMultimodal Perception

XPlainVerse evaluates deepfake detection and explanation quality, pairing real images with forgeries from twelve models and providing technical and simplified explanations, with metrics EntityScore and EvidenceScore for reasoning fidelity.

Released
2026-07-03
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks focus on classification accuracy, not explanation grounding. XPlainVerse aims to measure whether explanations are grounded in actual manipulations, which is key for trustworthy deployable detection.

Motivation

As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limiting reliability and user trust.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.