Benchmark Radar
AI BENCHMARK PROFILE

RNG-Bench

General AIMultimodal PerceptionInternLM

RNG-Bench evaluates multimodal LLMs in controllable non-Markov games, requiring reconstruction of past observations and acting on them, with two games: Matching Pairs and 3D Maze.

Released
2026-06-17
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the gap in evaluating models' abilities to remember and act on hidden state, which is critical for real-world deployments where observations are partial.

Motivation

Deploying multimodal foundation models as closed-loop policies increasingly requires conditioning actions on observations that are no longer visible.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.