Benchmark Radar
AI BENCHMARK PROFILE

GEB-Bench

General AIMultimodal PerceptionMathematics & Formal Sciences

GEB-Bench evaluates models on abstract structural motifs (e.g., self-reference, strange loops) presented in multiple modalities (natural scenes, folk stories, math theorems, programmatic skeletons) with tasks probing cross-modal mapping.

Released
2026-08-04
Readiness
Paper only
Primary field
General AI

Why it matters

The benchmark aims to measure abstraction and cross-modal transfer, which are foundational for general intelligence but not addressed by current benchmarks.

Motivation

Can a model look at a river delta and a lightning bolt and see that they share a structure?

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.