SQBench
SQBench v1.0 evaluates language-model agents on 220 production-oriented tasks organized into L1 atomic capabilities, L2 composite skills, and L3 business scenarios. Tasks require processing input assets, using tools, and producing a specified deliverable. Scoring computes Completion, Risk Penalty, and Performance from a 10D Risk Matrix; Strict Pass requires Completion=1 and Risk Penalty=0.
- Released
- 2026-07-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
The benchmark targets delivery under domain constraints, a shared weakness in current models. Its scoring separates functional completion from risk, which could support decisions about agent deployment in production workflows. However, without public access to tasks and full results, its practical value is limited for external comparison.
Motivation
Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.