AI BENCHMARK PROFILE
QuoteBench
Benchmark of 56 one-shot tasks from 14 incident-derived families that measures exact final-state outcomes when LLM coding agents issue Bash commands through serializing/wrapping/reparsing transport, using an added unescaped parser.
- Released
- 2026-08-13
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Reveals that matched execution scores can hide command-path failures, prompting evaluation of deployment configurations rather than treating model scores as intrinsic.
Motivation
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.