Benchmark Radar
AI BENCHMARK PROFILE

VendorBench-100

General AIMultimodal PerceptionSharayu Deshmukh

Evaluates deepfake image detectors across three paradigms—commercial APIs, vision LLMs, and open-source detectors—on a fixed 100-image adversarial corpus. Uses a unified output schema and scores primarily by Matthews correlation coefficient with ROC-AUC.

Released
2026-07-07
Readiness
Runnable
Primary field
General AI

Why it matters

Provides a common ground for comparing disparate detector types, addressing the lack of unified evaluation. Identifies metric correlation and calibration issues that matter for real-world deployment decisions.

Motivation

Deepfake image detection is served by three fundamentally different paradigms - commercial APIs, zero-shot vision-language models (LLMs), and open-source detectors - that are rarely evaluated under a common protocol, making direct comparison difficult.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.