Benchmark Radar
AI BENCHMARK PROFILE

StemBind

General AIMultimodal Perception

StemBind is a diagnostic benchmark with shared-stem questions to attribute failures in abstract visual reasoning. It includes perception, rule, and full tasks with stage annotations but no official code or dataset release.

Released
2026-05-29
Readiness
Inspectable
Primary field
General AI

Why it matters

It aims to localize reasoning failures to specific sub-steps, potentially guiding improvements in multimodal models, but lacks a public path for independent verification.

Motivation

Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can describe what it sees and name the underlying pattern, yet still fail to choose the matching candidate.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.