Benchmark Radar
AI BENCHMARK PROFILE

QuoteBench

General AIKnowledge & ReasoningQuoteBench team

Benchmark of 56 one-shot tasks from 14 incident-derived families that measures exact final-state outcomes when LLM coding agents issue Bash commands through serializing/wrapping/reparsing transport, using an added unescaped parser.

Released
2026-08-13
Readiness
Paper only
Primary field
General AI

Why it matters

Reveals that matched execution scores can hide command-path failures, prompting evaluation of deployment configurations rather than treating model scores as intrinsic.

Motivation

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.