Benchmark Radar
AI BENCHMARK PROFILE

GameXpert-Bench

General AIKnowledge & Reasoning

Evaluates coding agents across three game development lifecycle tracks: generation, bug diagnosis and repair, and multi-turn optimization, using live interaction, behavioral tests, and product criteria.

Released
2026-08-22
Readiness
Paper only
Primary field
General AI

Why it matters

Provides a comprehensive benchmark for game development with coding agents, measuring both product quality and process capabilities across the full lifecycle.

Motivation

Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.