Benchmark Radar
AI BENCHMARK PROFILE

GPTNT

General AIMultimodal PerceptionGPTNT Project

GPTNT evaluates multimodal agents on real-time collaborative bomb defusal in the game Keep Talking and Nobody Explodes, requiring asynchronous communication under time pressure and information asymmetry. Success is measured by defusing procedurally generated bombs.

Released
2026-06-26
Readiness
Inspectable
Primary field
General AI

Why it matters

Current benchmarks isolate collaboration components; GPTNT captures time pressure, information asymmetry, and imperfect communication together, providing a realistic test for multimodal systems that current evaluations leave unmeasured.

Motivation

Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.