Benchmark Radar
AI BENCHMARK PROFILE

RuleMaze

General AIMultimodal PerceptionOceanFlowLab

Benchmark for rule-compliant visual spatial planning in multimodal LLMs, requiring maze navigation under natural-language rules with automated rule generation and validation.

Released
2026-08-20
Readiness
Runnable
Primary field
General AI

Why it matters

Fills a gap in evaluating MLLMs on joint visual perception, rule interpretation, and constrained action planning, with scalable rule construction and a public leaderboard.

Motivation

Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.