Benchmark Radar
AI BENCHMARK PROFILE

DiG-bench

General AIKnowledge & ReasoningDiG-bench Team

Evaluates AI agents on discovery of unknown transformation rules across 70 independent games with seven difficulty tiers, requiring experimentation and rule generalization.

Released
2026-08-12
Readiness
Paper only
Primary field
General AI

Why it matters

Fills the gap in benchmarks for scientific discovery in controlled environments, testing the capacity to formulate novel generalizations through interaction.

Motivation

Discovery---formulating novel generalizations---is a central part of the scientific process.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.