Benchmark Radar
AI BENCHMARK PROFILE

Inverse Turing Bench

General AIKnowledge & Reasoning

The benchmark evaluates language models on distinguishing human-only vs. human-AI multi-turn dialogues, using paired transcripts and accuracy as the metric.

Released
2026-06-20
Readiness
Paper only
Primary field
General AI

Why it matters

It addresses the practical need for reliable human-AI differentiation in online spaces, with implications for trust and safety. The benchmark may help compare detection approaches, though its current form primarily supports the paper's findings.

Motivation

As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.