Benchmark Radar
AI BENCHMARK PROFILE

PatternEval

General AIMultimodal Perception

PatternEval is a diagnostic benchmark with 2,415 multimodal prompts testing four response-pattern failures: chain-of-thought leakage, repetition, contradiction, and performative reasoning.

Released
2026-08-13
Readiness
Paper only
Primary field
General AI

Why it matters

Evaluates response-pattern alignment between thinking and non-thinking modes in hybrid-thinking MLLMs, addressing user-facing quality beyond accuracy.

Motivation

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.