HarnessOpt-Bench
Evaluates LLMs optimizing a target agent's harness (prompts, tools, control flow) under budgeted, stochastic evaluation. Scoring is normalized gain over seed on a held-out test partition, with trusted execution environment enforcing evaluation boundary.
- Released
- 2026-08-06
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the gap in measuring automated harness optimization, a discriminative capability needed for improving agentic LLM systems. Provides a protocol for comparing optimizer models under varying harnesses and tasks.
Motivation
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.