AI BENCHMARK PROFILE
PatternEval
PatternEval is a diagnostic benchmark with 2,415 multimodal prompts testing four response-pattern failures: chain-of-thought leakage, repetition, contradiction, and performative reasoning.
- Released
- 2026-08-13
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Evaluates response-pattern alignment between thinking and non-thinking modes in hybrid-thinking MLLMs, addressing user-facing quality beyond accuracy.
Motivation
Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.