Benchmark Radar
AI BENCHMARK PROFILE

SteerBench-Work

Health & Life SciencesFinance & EconomicsKnowledge & Reasoning

SteerBench-Work evaluates agent steering decisions at action boundaries in workplace scenarios across seven domains, with 106 incident-anchored scenarios and evidence-reversed mirrors, scored on correct proceed/hold boundaries.

Released
2026-08-12
Readiness
Paper only
Primary field
Health & Life Sciences

Why it matters

Addresses the critical pre-commit decision in long-running agents, where a single step can have significant consequences. The benchmark reveals that models tend to over-refuse authorized actions, providing valuable insight for calibrating agent behavior to avoid both unsafe actions and unnecessary delays.

Motivation

Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.