AI BENCHMARK PROFILE
DiG-bench
Evaluates AI agents on discovery of unknown transformation rules across 70 independent games with seven difficulty tiers, requiring experimentation and rule generalization.
- Released
- 2026-08-12
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Fills the gap in benchmarks for scientific discovery in controlled environments, testing the capacity to formulate novel generalizations through interaction.
Motivation
Discovery---formulating novel generalizations---is a central part of the scientific process.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.