Benchmark Radar
AI BENCHMARK PROFILE

RevengeBench

General AIKnowledge & Reasoning

RevengeBench is a benchmark for recovering code-space policies from behavioral traces. It includes 75 LLM-generated policies across five game environments, where a learner designs behavioral probes and submits executable hypotheses, evaluated using continuous action-distance metrics.

Released
2026-06-24
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the inverse problem of inferring hidden decision programs from observations, relevant to opponent modeling and policy interpretability. It provides a tractable testbed for studying how controlled experiments improve code-space recovery.

Motivation

For most of scientific history, researchers studying behavior could only infer hidden mechanisms from outward actions: an inverse problem that becomes more tractable when observation is augmented by targeted intervention.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.