AI BENCHMARK PROFILE
GPTNT
GPTNT evaluates multimodal agents on real-time collaborative bomb defusal in the game Keep Talking and Nobody Explodes, requiring asynchronous communication under time pressure and information asymmetry. Success is measured by defusing procedurally generated bombs.
- Released
- 2026-06-26
- Readiness
- Inspectable
- Primary field
- General AI
Why it matters
Current benchmarks isolate collaboration components; GPTNT captures time pressure, information asymmetry, and imperfect communication together, providing a realistic test for multimodal systems that current evaluations leave unmeasured.
Motivation
Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.