Benchmark Radar
AI BENCHMARK PROFILE

SIMMER

General AIKnowledge & Reasoning

SIMMER evaluates latent failures in LLM-generated plans for kitchen-domain tasks using a curated symbolic world model with 77 actions and 262 objects, scoring error-free plans and latent hazard detection.

Released
2026-06-12
Readiness
Paper only
Primary field
General AI

Why it matters

Existing plan benchmarks miss failures that don't immediately halt execution but compromise goals; SIMMER provides metrics for irreversible latent failures, important for safe deployment of LLM planners.

Motivation

Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.