AI BENCHMARK PROFILE
PTCG-Bench
PTCG-Bench evaluates LLM agents on the Pokémon Trading Card Game at two levels: single-environment decision-making and self-evolution through accumulated experience, with modular harness ablation to separate agent performance from harness design.
- Released
- 2026-05-28
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Agent benchmarks often miss strategic and evolving decision-making. PTCG-Bench provides a realistic interactive environment to assess self-evolution and harness sensitivity, which are critical for deploying autonomous agents.
Motivation
Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.