Benchmark Radar
AI BENCHMARK PROFILE

PTCG-Bench

General AIAgents

PTCG-Bench evaluates LLM agents on the Pokémon Trading Card Game at two levels: single-environment decision-making and self-evolution through accumulated experience, with modular harness ablation to separate agent performance from harness design.

Released
2026-05-28
Readiness
Paper only
Primary field
General AI

Why it matters

Agent benchmarks often miss strategic and evolving decision-making. PTCG-Bench provides a realistic interactive environment to assess self-evolution and harness sensitivity, which are critical for deploying autonomous agents.

Motivation

Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.