AI BENCHMARK PROFILE
SIMMER
SIMMER evaluates latent failures in LLM-generated plans for kitchen-domain tasks using a curated symbolic world model with 77 actions and 262 objects, scoring error-free plans and latent hazard detection.
- Released
- 2026-06-12
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing plan benchmarks miss failures that don't immediately halt execution but compromise goals; SIMMER provides metrics for irreversible latent failures, important for safe deployment of LLM planners.
Motivation
Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.