Benchmark Radar
AI BENCHMARK PROFILE

SQBench

General AIKnowledge & Reasoning

SQBench v1.0 evaluates language-model agents on 220 production-oriented tasks organized into L1 atomic capabilities, L2 composite skills, and L3 business scenarios. Tasks require processing input assets, using tools, and producing a specified deliverable. Scoring computes Completion, Risk Penalty, and Performance from a 10D Risk Matrix; Strict Pass requires Completion=1 and Risk Penalty=0.

Released
2026-07-25
Readiness
Runnable
Primary field
General AI

Why it matters

The benchmark targets delivery under domain constraints, a shared weakness in current models. Its scoring separates functional completion from risk, which could support decisions about agent deployment in production workflows. However, without public access to tasks and full results, its practical value is limited for external comparison.

Motivation

Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.