Benchmark Radar
AI BENCHMARK PROFILE

AgentAbstain

General AIKnowledge & Reasoning

AgentAbstain is a paired-task benchmark for evaluating LLM agents' ability to abstain from acting in scenarios such as ambiguity, conflicting constraints, or tool failures. It includes 263 paired tasks across 42 sandbox environments, with a proposed pipeline for generating fresh task instances.

Released
2026-07-11
Readiness
Paper only
Primary field
General AI

Why it matters

Agent abstention is critical for safe deployment, yet existing evaluations focus on task success. This benchmark targets the gap in measuring calibrated abstention, highlighting that abstention capability is independent of general task-solving ability.

Motivation

Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.