AI BENCHMARK PROFILE
GameXpert-Bench
Evaluates coding agents across three game development lifecycle tracks: generation, bug diagnosis and repair, and multi-turn optimization, using live interaction, behavioral tests, and product criteria.
- Released
- 2026-08-22
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Provides a comprehensive benchmark for game development with coding agents, measuring both product quality and process capabilities across the full lifecycle.
Motivation
Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.